CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon-Aware Embeddings with Amino Acid-Aware Embeddings and Wavelet-based Kolmogorov–Arnold Network

Author:

Pratyush PawelORCID,Carrier Callen,Pokharel Suresh,Ismail Hamid D.,Chaudhari Meenal,KC Dukka B.

Abstract

AbstractThe mapping from codon to amino acid is surjective due to the high degeneracy of the codon alphabet, suggesting that codon space might harbor higher information content. Embeddings from the codon language model have recently demonstrated success in various downstream tasks. However, predictive models for phosphorylation sites, arguably the most studied Post-Translational Modification (PTM), and PTM sites in general, have predominantly relied on amino acid-level representations. This work introduces a novel approach for prediction of phosphorylation sites by incorporating codon-level information through embeddings from a recently developed codon language model trained exclusively on protein-coding DNA sequences. Protein sequences are first meticulously mapped to reliable coding sequences and encoded using this encoder to generate codon-aware embeddings. These embeddings are then integrated with amino acid-aware embeddings obtained from a protein language model through an early fusion strategy. Subsequently, a window-level representation of the site of interest is formed from the fused embeddings within a defined window frame. A ConvBiGRU network extracts features capturing spatiotemporal correlations between proximal residues within the window, followed by a Kolmogorov-Arnold Network (KAN) based on the Derivative of Gaussian (DoG) wavelet transform function to produce the prediction inference for the site. We dub the overall model integrating these elements as CaLMPhosKAN. On independent testing with Serine-Threonine (combined) and Tyrosine test sets, CaLMPhosKAN outperforms existing approaches. Furthermore, we demonstrate the model’s effectiveness in predicting sites within intrinsically disordered regions of proteins. Overall, CaLMPhosKAN emerges as a robust predictor of general phosphosites in proteins. CaLMPhosKAN will be released publicly soon.

Publisher

Cold Spring Harbor Laboratory

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3