Institute of Computing Technology, Chinese Academy IR
Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure prediction | |
Ju, Fusong1,2; Zhu, Jianwei3; Zhang, Qi1,2; Wei, Guozheng1,2; Sun, Shiwei1,2,4; Zheng, Wei-Mou2,5; Bu, Dongbo1,2,4 | |
2022-02-15 | |
发表期刊 | BIOINFORMATICS |
ISSN | 1367-4803 |
卷号 | 38期号:4页码:990-996 |
摘要 | Motivation: Accurate prediction of protein structure relies heavily on exploiting multiple sequence alignment (MSA) for residue mutations and correlations as this information specifies protein tertiary structure. The widely used prediction approaches usually transform MSA into inter-mediate models, say position-specific scoring matrix or profile hidden Markov model. These inter-mediate models, however, cannot fully represent residue mutations and correlations carried by MSA; hence, an effective way to directly exploit MSAs is highly desirable. Results: Here, we report a novel sequence set network (called Seq-SetNet) to directly and effectively exploit MSA for protein structure prediction. Seq-SetNet uses an `encoding and aggregation' strategy that consists of two key elements: (i) an encoding module that takes a component homologue in MSA as input, and encodes residue mutations and correlations into context-specific features for each residue; and (ii) an aggregation module to aggregate the features extracted from all component homologues, which are further transformed into structural properties for residues of the query protein. As Seq-SetNet encodes each homologue protein individually, it could consider both insertions and deletions, as well as long-distance correlations among residues, thus representing more information than the inter-mediate models. Moreover, the encoding module automatically learns effective features and thus avoids manual feature engineering. Using symmetric aggregation functions, Seq-SetNet processes the homologue proteins as a sequence set, making its prediction results invariable to the order of these proteins. On popular benchmark sets, we demonstrated the successful application of Seq-SetNet to predict secondary structure and torsion angles of residues with improved accuracy and efficiency. |
DOI | 10.1093/bioinformatics/btab777 |
收录类别 | SCI |
语种 | 英语 |
资助项目 | National Key Research and Development Program of China[2020YFA0907000] ; National Natural Science Foundation of China[62072435] ; National Natural Science Foundation of China[31770775] ; National Natural Science Foundation of China[82130055] |
WOS研究方向 | Biochemistry & Molecular Biology ; Biotechnology & Applied Microbiology ; Computer Science ; Mathematical & Computational Biology ; Mathematics |
WOS类目 | Biochemical Research Methods ; Biotechnology & Applied Microbiology ; Computer Science, Interdisciplinary Applications ; Mathematical & Computational Biology ; Statistics & Probability |
WOS记录号 | WOS:000747962400015 |
出版者 | OXFORD UNIV PRESS |
引用统计 | |
文献类型 | 期刊论文 |
条目标识符 | http://119.78.100.204/handle/2XEOYT63/19003 |
专题 | 中国科学院计算技术研究所期刊论文_英文 |
通讯作者 | Bu, Dongbo |
作者单位 | 1.Chinese Acad Sci, Inst Comp Technol, Key Lab Intelligent Informat Proc, Beijing 100190, Peoples R China 2.Univ Chinese Acad Sci, Beijing 100049, Peoples R China 3.Microsoft Res Asia, Beijing 100080, Peoples R China 4.Zhongke Big Data Acad, Zhengzhou 450046, Henan, Peoples R China 5.Chinese Acad Sci, Inst Theoret Phys, Beijing 100190, Peoples R China |
推荐引用方式 GB/T 7714 | Ju, Fusong,Zhu, Jianwei,Zhang, Qi,et al. Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure prediction[J]. BIOINFORMATICS,2022,38(4):990-996. |
APA | Ju, Fusong.,Zhu, Jianwei.,Zhang, Qi.,Wei, Guozheng.,Sun, Shiwei.,...&Bu, Dongbo.(2022).Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure prediction.BIOINFORMATICS,38(4),990-996. |
MLA | Ju, Fusong,et al."Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure prediction".BIOINFORMATICS 38.4(2022):990-996. |
条目包含的文件 | 条目无相关文件。 |
除非特别说明,本系统中所有内容都受版权保护,并保留所有权利。
修改评论