A method for multiple-sequence-alignment-free protein structure prediction using a protein language model-Reference-Cited by-同舟云学术

A method for multiple-sequence-alignment-free protein structure prediction using a protein language model

Published:2023-10-09 Issue:10 Volume:5 Page:1087-1096
ISSN:2522-5839
Container-title:Nature Machine Intelligence
language:en
Short-container-title:Nat Mach Intell

Author:

Fang Xiaomin,Wang Fan^ORCID,Liu Lihang^ORCID,He Jingzhou,Lin Dayong,Xiang Yingfei^ORCID,Zhu Kunrui,Zhang Xiaonan,Wu Hua,Li Hui,Song Le^ORCID

Abstract

AbstractProtein structure prediction pipelines based on artificial intelligence, such as AlphaFold2, have achieved near-experimental accuracy. These advanced pipelines mainly rely on multiple sequence alignments (MSAs) as inputs to learn the co-evolution information from the homologous sequences. Nonetheless, searching MSAs from protein databases is time consuming, usually taking tens of minutes. Consequently, we attempt to explore the limits of fast protein structure prediction by using only primary structures of proteins. Our proposed method, HelixFold-Single, combines a large-scale protein language model with the superior geometric learning capability of AlphaFold2. HelixFold-Single first pre-trains a large-scale protein language model with thousands of millions of primary structures utilizing the self-supervised learning paradigm, which will be used as an alternative to MSAs for learning the co-evolution information. Then, by combining the pre-trained protein language model and the essential components of AlphaFold2, we obtain an end-to-end differentiable model to predict the three-dimensional coordinates of atoms from only the primary structure. HelixFold-Single is validated on datasets CASP14 and CAMEO, achieving competitive accuracy with the MSA-based methods on targets with large homologous families. Furthermore, HelixFold-Single consumes much less time than the mainstream pipelines for protein structure prediction, demonstrating its potential in tasks requiring many predictions.

Publisher

Springer Science and Business Media LLC

Subject

Artificial Intelligence,Computer Networks and Communications,Computer Vision and Pattern Recognition,Human-Computer Interaction,Software

Link

https://www.nature.com/articles/s42256-023-00721-6.pdf

Reference34 articles.

1. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).

2. Moult, J. A decade of CASP: progress, bottlenecks and prognosis in protein structure prediction. Curr. Opin. Struct. Biol. 15, 285–289 (2005).

3. Petroni, F. et al. Language models as knowledge bases? In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) https://doi.org/10.18653/v1/D19-1250 (ACL, 2019).

4. Vaswani, A. et al. Attention is all you need. In NIPS'17: Proc. 31st International Conference on Neural Information Processing Systems Vol. 30 (eds von Luxburg, U. et al.) 6000–6010 (Curran, 2017).

5. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (eds Burstein, J. et al.) 4171–4186 (Association for Computational Linguistics, 2019).

Cited by 32 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The characteristic structural and functional dynamics of P. falciparum DHFR binding with pyrimidine chemotypes implicate malaria therapy design;Chemical Physics Impact;2024-12

2. Structural and functional prediction, evaluation, and validation in the post-sequencing era;Computational and Structural Biotechnology Journal;2024-12

3. Improving prediction performance of general protein language model by domain-adaptive pretraining on DNA-binding protein;Nature Communications;2024-09-07

4. AI-accelerated therapeutic antibody development: practical insights;Frontiers in Drug Discovery;2024-09-03

5. Unveiling the evolution of policies for enhancing protein structure predictions: A comprehensive analysis;Computers in Biology and Medicine;2024-09