Predicting hotspots for disease-causing single nucleotide variants using sequences-based coevolution, network analysis, and machine learning-Reference-Cited by-同舟云学术

Predicting hotspots for disease-causing single nucleotide variants using sequences-based coevolution, network analysis, and machine learning

Published:2024-05-14 Issue:5 Volume:19 Page:e0302504
ISSN:1932-6203
Container-title:PLOS ONE
language:en
Short-container-title:PLoS ONE

Author:

Zheng Wenjun^ORCID

Abstract

To enable personalized medicine, it is important yet highly challenging to accurately predict disease-causing mutations in target proteins at high throughput. Previous computational methods have been developed using evolutionary information in combination with various biochemical and structural features of protein residues to discriminate neutral vs. deleterious mutations. However, the power of these methods is often limited because they either assume known protein structures or treat residues independently without fully considering their interactions. To address the above limitations, we build upon recent progress in machine learning, network analysis, and protein language models, and develop a sequences-based variant site prediction workflow based on the protein residue contact networks: 1. We employ and integrate various methods of building protein residue networks using state-of-the-art coevolution analysis tools (RaptorX, DeepMetaPSICOV, and SPOT-Contact) powered by deep learning. 2. We use machine learning algorithms (Random Forest, Gradient Boosting, and Extreme Gradient Boosting) to optimally combine 20 network centrality scores to jointly predict key residues as hot spots for disease mutations. 3. Using a dataset of 107 proteins rich in disease mutations, we rigorously evaluate the network scores individually and collectively (via machine learning). This work supports a promising strategy of combining an ensemble of network scores based on different coevolution analysis methods (and optionally predictive scores from other methods) via machine learning to predict hotspot sites of disease mutations, which will inform downstream applications of disease diagnosis and targeted drug design.

Funder

NIH

Publisher

Public Library of Science (PLoS)

Reference69 articles.

1. Highly accurate protein structure prediction with AlphaFold;J Jumper;Nature,2021

2. Accurate prediction of protein structures and interactions using a three-track neural network;M Baek;Science,2021

3. AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination;TC Terwilliger;Nat Methods,2023

4. Has DeepMind’s AlphaFold solved the protein folding problem?;A Al-Janabi;Biotechniques,2022

5. Unraveling protein’s structural dynamics: from configurational dynamics to ensemble switching guides functional mesoscale assemblies;E Medina;Curr Opin Struct Biol,2021