Linc2function: A Comprehensive Pipeline and Webserver for Long Non-Coding RNA (lncRNA) Identification and Functional Predictions Using Deep Learning Approaches-Reference-Cited by-同舟云学术

Linc2function: A Comprehensive Pipeline and Webserver for Long Non-Coding RNA (lncRNA) Identification and Functional Predictions Using Deep Learning Approaches

Published:2023-09-15 Issue:3 Volume:7 Page:22
ISSN:2075-4655
Container-title:Epigenomes
language:en
Short-container-title:Epigenomes

Author:

Ramakrishnaiah Yashpal¹²^ORCID,Morris Adam P.³,Dhaliwal Jasbir²,Philip Melcy¹^ORCID,Kuhlmann Levin⁴,Tyagi Sonika¹²

Affiliation:

1. Central Clinical School, Monash University, Melbourne, VIC 3000, Australia

2. School of Computing Technologies, Royal Melbourne Institute of Technology University, Melbourne, VIC 3000, Australia

3. Monash Data Futures Institute, Monash University, Clayton, VIC 3800, Australia

4. Faculty of Information Technology, Monash University, Clayton, VIC 3800, Australia

Abstract

Long non-coding RNAs (lncRNAs), comprising a significant portion of the human transcriptome, serve as vital regulators of cellular processes and potential disease biomarkers. However, the function of most lncRNAs remains unknown, and furthermore, existing approaches have focused on gene-level investigation. Our work emphasizes the importance of transcript-level annotation to uncover the roles of specific transcript isoforms. We propose that understanding the mechanisms of lncRNA in pathological processes requires solving their structural motifs and interactomes. A complete lncRNA annotation first involves discriminating them from their coding counterparts and then predicting their functional motifs and target bio-molecules. Current in silico methods mainly perform primary-sequence-based discrimination using a reference model, limiting their comprehensiveness and generalizability. We demonstrate that integrating secondary structure and interactome information, in addition to using transcript sequence, enables a comprehensive functional annotation. Annotating lncRNA for newly sequenced species is challenging due to inconsistencies in functional annotations, specialized computational techniques, limited accessibility to source code, and the shortcomings of reference-based methods for cross-species predictions. To address these challenges, we developed a pipeline for identifying and annotating transcript sequences at the isoform level. We demonstrate the effectiveness of the pipeline by comprehensively annotating the lncRNA associated with two specific disease groups. The source code of our pipeline is available under the MIT licensefor local use by researchers to make new predictions using the pre-trained models or to re-train models on new sequence datasets. Non-technical users can access the pipeline through a web server setup.

Funder

Monash University’s Australian Women in Research Acceleration

National Health and Medical Research Council

Publisher

MDPI AG

Subject

Health, Toxicology and Mutagenesis,Genetics,Biochemistry, Genetics and Molecular Biology (miscellaneous),Biochemistry

Link

https://www.mdpi.com/2075-4655/7/3/22/pdf

Reference46 articles.

1. Genome-wide computational identification and manual annotation of human long noncoding RNA genes;Jia;RNA,2010

2. Non-coding RNA;Mattick;Hum. Mol. Genet.,2006

3. Discovery and functional analysis of lncRNAs: Methodologies to investigate an uncharacterized transcriptome;Kashi;Biochim. Biophys. Acta (BBA) Gene Regul. Mech.,2016

4. van Bakel, H., Nislow, C., Blencowe, B.J., and Hughes, T.R. (2010). Most “Dark Matter” Transcripts Are Associated with Known Genes. PLoS Biol., 8.

5. Agrawal, S., Alam, T., Koido, M., Kulakovskiy, I.V., Severin, J., Abugessaisa, I., Buyan, A., Dostie, J., Itoh, M., and Kondo, N. (2021). Functional annotation of human long noncoding RNAs using chromatin conformation data. bioRxiv, bioRxiv:2021.01.13.426305.

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Automated Navigation of the lncRNA Transcriptome: A comprehensive SnakeMake based computational Pipeline for robust Identification of lncRNAs and their putative targets;2024-08-19

2. Long Intergenic Non-Coding RNAs of Human Chromosome 18: Focus on Cancers;Biomedicines;2024-02-28