SolXplain: An Explainable Sequence-Based Protein Solubility Predictor-Reference-Cited by-同舟云学术

SolXplain: An Explainable Sequence-Based Protein Solubility Predictor

Published:2019-05-27 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Mall Raghvendra

Abstract

AbstractMotivationProtein solubility is a property associated with protein expression and is a critical determinant of the manufacturability of therapeutic proteins. It is thus imperative to design accurate in-silico sequence-based solubility predictors.MethodsIn this study, we propose SolXplain, an extreme gradient boosting machine based protein solubility predictor which achieves state-of-the-art performance using physio-chemical, sequence and novel structure derived features from protein sequences. Moreover, SolXplain has a unique attribute that it can provide explanation for the predicted class label for each test protein based on its corresponding feature values using SHapley Additive exPlanations (SHAP) method.ResultsBased on an independent test set, SolXplain outperformed other sequence-based methods by at least 2% in accuracy and 2% in Matthew’s correlation coefficient, with an overall accuracy of 78% and Matthew’s correlation coefficient of 0.56. Additionally, for fractions of exposed residues (FER) at various residual solvent accessibility (RSA) cutoffs, we observed higher fractions to associate positively with protein solubility, and tripeptide stretches that contain one isoleucine and one or more histidines, to associate negatively with solubility. The improved prediction accuracy of SolXplain enables it to predict protein solubility with greater consistency and screen for sequences with enhanced manufacturability.

Publisher

Cold Spring Harbor Laboratory

Reference42 articles.

1. Understanding the relationship between the primary structure of proteins and its propensity to be soluble on overexpression inEscherichia coli

2. SOLpro: accurate sequence-based prediction of protein solubility

3. Predicting the Solubility of Recombinant Proteins in Escherichia coli

4. New fusion protein systems designed to give soluble expression inEscherichia coli

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. PyPEF—An Integrated Framework for Data-Driven Protein Engineering;Journal of Chemical Information and Modeling;2021-07-14

2. Structure-aware Protein Solubility Prediction From Sequence Through Graph Convolutional Network And Predicted Contact Map;2020-06-25