Author:
Chandra Omkar,Sharma Madhu,Pandey Neetesh,Jha Indra Prakash,Mishra Shreya,Kong Say Li,Kumar Vibhor
Abstract
AbstractThe number of annotated genes in the human genome has increased tremendously, and understanding their biological role is challenging through experimental methods alone. There is a need for a computational approach to infer the function of genes, particularly for non-coding RNAs, with reliable explainability. We have utilized genomic features that are present across both coding and non-coding genes like transcription factor (TF) binding pattern, histone modifications, and DNase hypersensitivity profiles to predict ontology-based functions of genes. Our approach for gene function prediction (GFPred) made reliable predictions (>90% balanced accuracy) for 486 gene-sets. Further analysis revealed that predictability using only TF-binding patterns at promoters is also high, and it paved the way for studying the effect of their combinatorics. The predicted associations between functions and genes were validated for their reliability using PubMed abstract mining. Clustering functions based on shared top predictive TFs revealed many latent groups of gene-sets involved in common major biological processes. Available CRISPR screens also supported the inferred association of genes with the major biological processes of latent groups of gene-sets. For the explainability of our approach, we also made more insights into the effect of combinatorics of TF binding (especially TF-pairs) on association with biological functions.
Publisher
Cold Spring Harbor Laboratory