基于傅里叶功率谱的H1N1病毒血凝素蛋白质序列的比较分析
Comparison and Analysis of H1N1 Hemagglutinin Protein Sequences Based on Fourier Power Spectrum
DOI: 10.12677/HJCB.2018.81003, PDF,    科研立项经费支持
作者: 王华*, 白凤兰, 刘立伟:大连交通大学理学院,辽宁 大连
关键词: 蛋白质序列傅里叶变换功率谱聚类Protein Sequence Fourier Transform Power Spectrum Clustering
摘要: 基于经典的HP模型,将不同特征下的H1N1病毒血凝素蛋白质序列转换为数字序列并且用离散傅里叶变换求出相应序列的功率谱。根据这些功率谱建立数学矩函数,并将数字序列转换为多维的矩向量,得到蛋白质序列对应的特征向量。再利用特征向量之间的中间距离对蛋白质序列进行聚类比较分析,得到了较好的结果。这一方法将不同长度的蛋白质序列通过功率谱和力矩将其转化为相同维数的向量,使我们更加容易比较分析生物序列。
Abstract: Based on the classical HP model, the H1N1 hemagglutinin protein sequence under different characteristics was converted into a digital sequence and the power spectrum of the corresponding sequence was calculated using a discrete Fourier transform. According to these power spectra, a mathematical moment function is established, and the digital sequence is converted into a multi-dimensional moment vector to obtain the corresponding feature vector of the protein sequence. Then using the middle distances between the feature vectors to compare and analyze the protein sequences, a good result was obtained. This method converts protein sequences of different lengths through power spectrum and moments into vectors of the same dimension, which makes it easier for us to compare and analyze biological sequences.
文章引用:王华, 白凤兰, 刘立伟. 基于傅里叶功率谱的H1N1病毒血凝素蛋白质序列的比较分析[J]. 计算生物学, 2018, 8(1): 15-23. https://doi.org/10.12677/HJCB.2018.81003

参考文献

[1] Katoh, K., Misawa, K.K.-I. and Miyata, T. (2002) Mafft: A Novel Method for Rapid Multiple Sequence Alignment Based on Fast Fourier Transform. Nucleic Acids Research, 30, 3059-3066. [Google Scholar] [CrossRef] [PubMed]
[2] Edgar, R.C. (2004) Muscle: Multiple Sequence Alignment with High Accuracy and High Throughput. Nucleic Acids Research, 32, 1792-1797. [Google Scholar] [CrossRef] [PubMed]
[3] Larkin, M.A., Blacksshields, G., Brown, N., Chenna, R., McGettigan, P.A., McWilliam, H., Valentin, F., Wallace, I.M., Wilm, A., Lopez, R., et al. (2007) Clustal w and Clustal x Version 2.0. Bioinformatics, 23, 2947-2948. [Google Scholar] [CrossRef] [PubMed]
[4] Marra, M.A., Jones, S.J., Astell, C.R., Holt, R.A., Brooks-Wilson, A., Butterfield, Y.S., Khattra, J., Asano, J.K., Barber, S.A., Chan, S.Y., et al. (2003) The Genome Sequence of the Sars-Associated Coronavirus. Science, 300, 1399-1404. [Google Scholar] [CrossRef] [PubMed]
[5] Vinga, S. and Almeida, J. (2003) Alignment-Free Sequence Comparison—A Review. Bioinformatics, 19, 513-523. [Google Scholar] [CrossRef] [PubMed]
[6] Yau, S.S.-T., Yu, C. and He, R. (2008) A Protein Map and Its Application. DNA and Cell Biology, 27, 241-250. [Google Scholar] [CrossRef] [PubMed]
[7] Yu, C., Deng, M. and Yau, S.S.-T. (2011) DNA Sequence Comparison by a Novel Probabilistic Method. Information Sciences, 18, 1484-1492. [Google Scholar] [CrossRef
[8] Yu, C., Hernandez, T., Zheng, H., Yau, S.-C., Huang, H.-H., He, R.L., Yang, J. and Yau, S.S.-T. (2013) Real Time Classification of Viruses in 12 Dimensions. PloS One, 8, e64328. [Google Scholar] [CrossRef] [PubMed]
[9] Pandit, A. and Sinha, S. (2010) Using Genomic Signatures for HIV-1 Sub-Typing. BMC Bioinformatics, 11, S26. [Google Scholar] [CrossRef
[10] Blaisdell, B.E. (1989) Average Values of a Dissimilarity Measure Not Requiring Sequence Alignment for a Computer-Generated Model System. Journal of Molecular Evolution, 29, 538-547. [Google Scholar] [CrossRef
[11] Anastassiou, D. (2000) Frequency-Domain Analysis of Biomolecular Sequences. Bioinformatics, 16, 1073-1081. [Google Scholar] [CrossRef] [PubMed]
[12] Kotlar, D. and Lavner, Y. (2003) Gene Prediction by Spectral Ro-tation Measure: A New Method for Identifying Protein Coding Regions. Genome Research, 13, 1930-1937. [Google Scholar] [CrossRef] [PubMed]
[13] Fukushima, A., Ikemura, T., Kinouchi, M., Oshima, T., Kodo, Y., Mori, H. and Kanaya, S. (2002) Periodicity in Prokaryotic and Eukaryotic Genomes Identified by Power Spectrum Analysis. Gene, 300, 203-211. [Google Scholar] [CrossRef
[14] Yin, C. and Yau, S.S.-T. (2005) A Fourier Characteristic of Coding Sequences: Origins and a Non-Fourier Apporximation. Journal of Computational Biology, 12, 1153-1165. [Google Scholar] [CrossRef] [PubMed]
[15] Yin, C. and Yau, S.S.-T. (2007) Prediction of Protein Coding Regions by the 3-Case Periodicity Analysis of DNA Sequence. Journal of Theoretical Biology, 247, 687-694. [Google Scholar] [CrossRef] [PubMed]
[16] Steinbach, M., Karypis, G. and Kumar, V. (2002) A Comparison of Docu-ment Clustering Techniques. KDD Workshop on Text Mining, 1-20.
[17] Hall, L.O. (2013) Exploring Big Data with Scalable Soft Clustering. Springer, Berlin Heidelberg, 11-15. [Google Scholar] [CrossRef
[18] 谢佳新, 殷建华, 李淑华, 鹿文英, 韩一芳, 韩磊, 张宏伟, 曹广文. 2009年新型甲型H1N1流感病毒血凝素基因进化分析[J]. 第二军医大学学报, 2009, 30(6): 613-617.
[19] Zhao, B., Duan, V. and Yau, S.S.T. (2011) A Novel Clustering Method via Nucleotid-Based Fourier Power Spectrum Analysis. Journal of Theoretical Biology, 279, 83-89. [Google Scholar] [CrossRef] [PubMed]
[20] 赵剑, 阮越, 王嘉松. 数学结构的蛋白质二维数字表达及其应用[J]. 数据采用与处理, 2013, 28(11): 770-776.
[21] 梁启浩, 李阳, 唐旭清. 基于功率谱的流感病毒蛋白质序列结构分析[J]. 病毒学报, 2017, 33(3): 313-319.
[22] Hoang, T., Yin, C., Zheng, H., Yu, C., He, R.L. and Yan, S.S.T. (2015) A New Method to Cluster DNA Sequences using Fourier Power Spectrum. Journal of Theoretical Biology, 372, 135-145. [Google Scholar] [CrossRef] [PubMed]
[23] 靳佩轩, 高洁. 流感病毒组成蛋白质序列的分析与预测[J]. 食品与生物技术学报, 2016, 35(4): 393-398.
[24] 李巍巍, 李阳, 唐旭清. 不同特征描述下H1N1病毒血凝素蛋白质序列的比较分析[J]. 生命科学研究, 2016, 20(2): 119-124.