TLCrys: Transfer Learning Based Method for Protein Crystallization Prediction-Reference-Cited by-同舟云学术

TLCrys: Transfer Learning Based Method for Protein Crystallization Prediction

Published:2022-01-16 Issue:2 Volume:23 Page:972
ISSN:1422-0067
Container-title:International Journal of Molecular Sciences
language:en
Short-container-title:IJMS

Author:

Jin Chen^ORCID,Shi Zhuangwei,Kang Chuanze,Lin Ken,Zhang Han

Abstract

X-ray diffraction technique is one of the most common methods of ascertaining protein structures, yet only 2–10% of proteins can produce diffraction-quality crystals. Several computational methods have been proposed so far to predict protein crystallization. Nevertheless, the current state-of-the-art computational methods are limited by the scarcity of experimental data. Thus, the prediction accuracy of existing models hasn’t reached the ideal level. To address the problems above, we propose a novel transfer-learning-based framework for protein crystallization prediction, named TLCrys. The framework proceeds in two steps: pre-training and fine-tuning. The pre-training step adopts attention mechanism to extract both global and local information of the protein sequences. The representation learned from the pre-training step is regarded as knowledge to be transferred and fine-tuned to enhance the performance of crystalization prediction. During pre-training, TLCrys adopts a multi-task learning method, which not only improves the learning ability of protein encoding, but also enhances the robustness and generalization of protein representation. The multi-head self-attention layer guarantees that different levels of the protein representation can be extracted by the fine-tuned step. During transfer learning, the fine-tuning strategy used by TLCrys improves the task-specialized learning ability of the network. Our method outperforms all previous predictors significantly in five crystallization stages of prediction. Furthermore, the proposed methodology can be well generalized to other protein sequence classification tasks.

Funder

National Natural Science Foundation of China

Publisher

MDPI AG

Subject

Inorganic Chemistry,Organic Chemistry,Physical and Theoretical Chemistry,Computer Science Applications,Spectroscopy,Molecular Biology,General Medicine,Catalysis

Link

https://www.mdpi.com/1422-0067/23/2/972/pdf

Reference32 articles.

1. The success of structural genomics

2. High Resolution NMR: Theory and Chemical Applications;Becker,1999

3. 15:30 STRUCTURAL ELUCIDATION OF DISC1 PATHWAY PROTEINS USING ELECTRON MICROSCOPY, CHEMICAL CROSS-LINKING AND MASS SPECTROSCOPY

4. Lessons from Structural Genomics

5. Structural Genomics, Round 2

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. High-Intensified Resemblance and Statistic-Restructured Alignment in Few-Shot Domain Adaptation for Industrial-Specialized Employment;IEEE Transactions on Consumer Electronics;2023-08

2. A Novel Classification Method for Antimicrobial Peptides Based on ProteinBERT;2023 42nd Chinese Control Conference (CCC);2023-07-24

3. Progress of the “Molecular Informatics” Section in 2022;International Journal of Molecular Sciences;2023-05-29

4. XGBoost and CNN-LSTM hybrid model with Attention-based stock prediction;2023 IEEE 3rd International Conference on Electronic Technology, Communication and Information (ICETCI);2023-05-26

5. Advances in protein solubility and thermodynamics: quantification, instrumentation, and perspectives;CrystEngComm;2023