Rapid protein evolution by few-shot learning with a protein language model-Reference-Cited by-同舟云学术

Rapid protein evolution by few-shot learning with a protein language model

Published:2024-07-18 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Jiang Kaiyi,Yan Zhaoqing,Bernardo Matteo Di,Sgrizzi Samantha R.,Villiger Lukas,Kayabolen Alisan,Kim Byungji,Carscadden Josephine K.,Hiraizumi Masahiro,Nishimasu Hiroshi,Gootenberg Jonathan S.^ORCID,Abudayyeh Omar O.^ORCID

Abstract

AbstractDirected evolution of proteins is critical for applications in basic biological research, therapeutics, diagnostics, and sustainability. However, directed evolution methods are labor intensive, cannot efficiently optimize over multiple protein properties, and are often trapped by local maxima.In silico-directed evolution methods incorporating protein language models (PLMs) have the potential to accelerate this engineering process, but current approaches fail to generalize across diverse protein families. We introduce EVOLVEpro, a few-shot active learning framework to rapidly improve protein activity using a combination of PLMs and protein activity predictors, achieving improved activity with as few as four rounds of evolution. EVOLVEpro substantially enhances the efficiency and effectiveness ofin silicoprotein evolution, surpassing current state-of-the-art methods and yielding proteins with up to 100-fold improvement of desired properties. We showcase EVOLVEpro for five proteins across three applications: T7 RNA polymerase for RNA production, a miniature CRISPR nuclease, a prime editor, and an integrase for genome editing, and a monoclonal antibody for epitope binding. These results demonstrate the advantages of few-shot active learning with small amounts of experimental data over zero-shot predictions. EVOLVEpro paves the way for broader applications of AI-guided protein engineering in biology and medicine.

Publisher

Cold Spring Harbor Laboratory

Reference54 articles.

1. Evolutionary-scale prediction of atomic-level protein structure with a language model

2. M. Heinzinger , K. Weissenow , J. G. Sanchez , A. Henkel , M. Mirdita , M. Steinegger , B. Rost , Bilingual Language Model for Protein Sequence and Structure, bioRxiv (2024)p. 2023.07.23.550085.

3. A. Elnaggar , H. Essam , W. Salah-Eldin , W. Moustafa , M. Elkerdawy , C. Rochereau , B. Rost , Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http://arxiv.org/abs/2301.06568.

4. ProteinBERT: a universal deep-learning model of protein sequence and function;Bioinformatics,2022

5. Protein language models-assisted optimization of a uracil-N-glycosylase variant enables programmable T-to-G and T-to-C base editing;Mol. Cell,2024

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Benchmarking text-integrated protein language model embeddings and embedding fusion on diverse downstream tasks;2024-08-26

2. Active Learning-Assisted Directed Evolution;2024-07-28