DeepSSPred: A Deep Learning Based Sulfenylation Site Predictor Via a Novel nSegmented Optimize Federated Feature Encoder-Reference-Cited by-同舟云学术

DeepSSPred: A Deep Learning Based Sulfenylation Site Predictor Via a Novel nSegmented Optimize Federated Feature Encoder

Published:2021-06-25 Issue:6 Volume:28 Page:708-721
ISSN:0929-8665
Container-title:Protein & Peptide Letters
language:en
Short-container-title:PPL

Author:

Khan Zaheer Ullah¹^ORCID,Pi Dechang¹^ORCID

Affiliation:

1. College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing, China

Abstract

Background: S-sulfenylation (S-sulphenylation, or sulfenic acid) proteins, are special kinds of post-translation modification, which plays an important role in various physiological and pathological processes such as cytokine signaling, transcriptional regulation, and apoptosis. Despite these aforementioned significances, and by complementing existing wet methods, several computational models have been developed for sulfenylation cysteine sites prediction. However, the performance of these models was not satisfactory due to inefficient feature schemes, severe imbalance issues, and lack of an intelligent learning engine. Objective: In this study, our motivation is to establish a strong and novel computational predictor for discrimination of sulfenylation and non-sulfenylation sites. Methods: In this study, we report an innovative bioinformatics feature encoding tool, named DeepSSPred, in which, resulting encoded features is obtained via nSegmented hybrid feature, and then the resampling technique called synthetic minority oversampling was employed to cope with the severe imbalance issue between SC-sites (minority class) and non-SC sites (majority class). State of the art 2D-Convolutional Neural Network was employed over rigorous 10-fold jackknife cross-validation technique for model validation and authentication. Results: Following the proposed framework, with a strong discrete presentation of feature space, machine learning engine, and unbiased presentation of the underline training data yielded into an excellent model that outperforms with all existing established studies. The proposed approach is 6% higher in terms of MCC from the first best. On an independent dataset, the existing first best study failed to provide sufficient details. The model obtained an increase of 7.5% in accuracy, 1.22% in Sn, 12.91% in Sp and 13.12% in MCC on the training data and12.13% of ACC, 27.25% in Sn, 2.25% in Sp, and 30.37% in MCC on an independent dataset in comparison with 2nd best method. These empirical analyses show the superlative performance of the proposed model over both training and Independent dataset in comparison with existing literature studies. Conclusion: In this research, we have developed a novel sequence-based automated predictor for SC-sites, called DeepSSPred. The empirical simulations outcomes with a training dataset and independent validation dataset have revealed the efficacy of the proposed theoretical model. The good performance of DeepSSPred is due to several reasons, such as novel discriminative feature encoding schemes, SMOTE technique, and careful construction of the prediction model through the tuned 2D-CNN classifier. We believe that our research work will provide a potential insight into a further prediction of S-sulfenylation characteristics and functionalities. Thus, we hope that our developed predictor will significantly helpful for large scale discrimination of unknown SC-sites in particular and designing new pharmaceutical drugs in general.

Publisher

Bentham Science Publishers Ltd.

Subject

Biochemistry,General Medicine,Structural Biology

Reference65 articles.

1. Voet D.; Voet J.G.; Pratt C.W.; Fundamentals of biochemistry: life at the molecular level 2013

2. Khoury G.A.; Baliban R.C.; Floudas C.A.; Proteome-wide post-translational modification statistics: frequency analysis and curation of the swiss-prot database. Sci Rep 2011,1,90

3. PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Res Hornbeck, P.V2011,40(D1),D261-D270

4. Mann M.; Jensen O.N.; Proteomic analysis of post-translational modifications. Nat Biotechnol 2003,21(3),255-261

5. Papin J.A.; Hunter T.; Palsson B.O.; Subramaniam S.; Reconstruction of cellular signalling networks and analysis of their properties. Nat Rev Mol Cell Biol 2005,6(2),99-111

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. DeepPhoPred: Accurate Deep Learning Model to Predict Microbial Phosphorylation;Proteins: Structure, Function, and Bioinformatics;2024-09-06

2. Glutathione kinetically outcompetes reactions between dimedone and a cyclic sulfenamide or physiological sulfenic acids;Free Radical Biology and Medicine;2023-11

3. Mini-review: Recent advances in post-translational modification site prediction based on deep learning;Computational and Structural Biotechnology Journal;2022

4. Prediction of lysine formylation sites using support vector machine based on the sample selection from majority classes and synthetic minority over-sampling techniques;Biochimie;2022-01