Masked Inverse Folding with Sequence Transfer for Protein Representation Learning-Reference-Cited by-同舟云学术

Masked Inverse Folding with Sequence Transfer for Protein Representation Learning

Published:2022-05-28 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Yang Kevin K.,Yeh Hugh,Zanichelli Niccolò

Abstract

AbstractSelf-supervised pretraining on protein sequences has led to state-of-the art performance on protein function and fitness prediction. However, sequence-only methods ignore the rich information contained in experimental and predicted protein structures. Meanwhile, inverse folding methods reconstruct a protein’s amino-acid sequence given its structure, but do not take advantage of sequences that do not have known structures. In this study, we train a masked inverse folding protein masked language model parameterized as a structured graph neural network. During pretraining, this model learns to reconstruct corrupted sequences conditioned on the backbone structure. We then show that using the outputs from a pretrained sequence-only protein masked language model as input to the inverse folding model further improves pretraining perplexity. We evaluate both of these models on downstream protein engineering tasks and analyze the effect of using information from experimental or predicted structures on performance.

Publisher

Cold Spring Harbor Laboratory

Reference74 articles.

1. The Rosetta all-atom energy function for macromolecular modeling and design;Journal of chemical theory and computation,2017

2. Unified rational protein engineering with sequence-based deep representation learning;Nat. Methods,2019

3. De novo protein design by deep network hallucination;Nature,2021

4. Tristan Bepler and Bonnie Berger . Learning protein sequence embeddings using information from structure. In International Conference on Learning Representations, 2019.

5. Nadav Brandes , Dan Ofer , Yam Peleg , Nadav Rappoport , and Michal Linial . ProteinBERT: A universal deep-learning model of protein sequence and function. bioRxiv, 2021.

Cited by 29 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. PETA: evaluating the impact of protein transfer learning with sub-word tokenization on downstream applications;Journal of Cheminformatics;2024-08-02

2. Generative artificial intelligence for de novo protein design;Current Opinion in Structural Biology;2024-06

3. GeoAB: Towards Realistic Antibody Design and Reliable Affinity Maturation;2024-05-17

4. NeuroFold: A Multimodal Approach to Generating Novel Protein Variantsin silico;2024-03-14

5. Protein Design Using Structure-Prediction Networks: AlphaFold and RoseTTAFold as Protein Structure Foundation Models;Cold Spring Harbor Perspectives in Biology;2024-03-04