Representational Learning from Healthy Multi-Tissue Human RNA-seq Data such that Latent Space Arithmetics Extracts Disease Modules-Reference-Cited by-同舟云学术

Representational Learning from Healthy Multi-Tissue Human RNA-seq Data such that Latent Space Arithmetics Extracts Disease Modules

Published:2023-10-05 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

de Weerd Hendrik A^ORCID,Guala Dimitri,Gustafsson Mika,Synnergren Jane,Tegnér Jesper,Lubovac-Pilav Zelmina,Magnusson Rasmus^ORCID

Abstract

1AbstractDeveloping computational analyses of transcriptomic data has dramatically improved our understanding of complex multifactorial diseases. However, such approaches are limited to small sample sets of disease-affected material, thus being sensitive to statistical biases and noise. Here, we ask if a variational autoencoder (VAE) trained on large groups of healthy, human RNA-seq data of multiple tissues can capture the fundamental healthy gene regulation system such that the learned representation generalizes to account for unseen disease changes. To this end, we trained a multi-scale representation to encode cellular processes ranging from cell types to genegene interactions. Importantly, we found that the learned healthy representations could predict unseen gene expression changes from 25 independent disease datasets. We extracted and decoded disease-specific signals from the VAE latent space to dissect this finding. Interestingly, the gene modules corresponding to this signal contained more disease-specific genes than the respective differential expression analysis in 20 of 25 cases. Finally, we matched genes related to the disease signals to known drug targets. We could extract sets of known and potential pharmaceutical candidates from this analysis and demonstrate the utility in three use cases. In summary, our study showcases how data-driven representation learning using a VAE as a foundational model allows an arithmetic deconstruction of the latent space such that biological insights enable the dissection of disease mechanisms and drug targets. Our model is available athttps://github.com/ddeweerd/VAE_Transcriptomics/.

Publisher

Cold Spring Harbor Laboratory

Reference47 articles.

1. Life's Complexity Pyramid

2. A DIseAse MOdule Detection (DIAMOnD) Algorithm Derived from a Systematic Analysis of Connectivity Patterns of Disease Proteins in the Human Interactome

3. A disease module in the interactome explains disease heterogeneity, drug response and captures novel pathways and genes in asthma

4. Community detection in networks: A user guide

5. Dynamic Response Genes in CD4+ T Cells Reveal a Network of Interactive Proteins that Classifies Disease Activity in Multiple Sclerosis

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Robust evaluation of deep learning-based representation methods for survival and gene essentiality prediction on bulk RNA-seq data;Scientific Reports;2024-07-24

2. Robust Evaluation of Deep Learning-based Representation Methods for Survival and Gene Essentiality Prediction on Bulk RNA-seq Data;2024-01-26