PLM‐T3SE: Accurate Prediction of Type III Secretion Effectors Using Protein Language Model Embeddings-Reference-Cited by-同舟云学术

PLM‐T3SE: Accurate Prediction of Type III Secretion Effectors Using Protein Language Model Embeddings

Published:2024-08-20 Issue: Volume: Page:
ISSN:0730-2312
Container-title:Journal of Cellular Biochemistry
language:en
Short-container-title:J of Cellular Biochemistry

Author:

Gao Mengru¹^ORCID,Song Chen¹,Liu Taigang¹

Affiliation:

1. College of Information Technology Shanghai Ocean University Shanghai China

Abstract

ABSTRACTThe Type III secretion effectors (T3SEs) are bacterial proteins synthesized by Gram‐negative pathogens and delivered into host cells via the Type III secretion system (T3SS). These effectors usually play a pivotal role in the interactions between bacteria and hosts. Hence, the precise identification of T3SEs aids researchers in exploring the pathogenic mechanisms of bacterial infections. Since the diversity and complexity of T3SE sequences often make traditional experimental methods time‐consuming, it is imperative to explore more efficient and convenient computational approaches for T3SE prediction. Inspired by the promising potential exhibited by pre‐trained language models in protein recognition tasks, we proposed a method called PLM‐T3SE that utilizes protein language models (PLMs) for effective recognition of T3SEs. First, we utilized PLM embeddings and evolutionary features from the position‐specific scoring matrix (PSSM) profiles to transform protein sequences into fixed‐length vectors for model training. Second, we employed the extreme gradient boosting (XGBoost) algorithm to rank these features based on their importance. Finally, a MLP neural network model was used to predict T3SEs based on the selected optimal feature set. Experimental results from the cross‐validation and independent test demonstrated that our model exhibited superior performance compared to the existing models. Specifically, our model achieved an accuracy of 98.1%, which is 1.8%–42.4% higher than the state‐of‐the‐art predictors based on the same independent data set test. These findings highlight the superiority of the PLM‐T3SE and the remarkable characterization ability of PLM embeddings for T3SE prediction.

Funder

National Natural Science Foundation of China

Publisher

Wiley

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/jcb.30642

Reference44 articles.

1. Protein delivery into eukaryotic cells by type III secretion machines

2. Type III Protein Secretion in Plant Pathogenic Bacteria

3. Functional domains and motifs of bacterial type III effector proteins and their roles in infection

4. Microbial genome-enabled insights into plant–microorganism interactions

5. Bacterial Type III Secretion Systems: Specialized Nanomachines for Protein Delivery into Target Cells