Rummagene: massive mining of gene sets from supporting materials of biomedical research publications-Reference-Cited by-同舟云学术

Rummagene: massive mining of gene sets from supporting materials of biomedical research publications

Published:2024-04-20 Issue:1 Volume:7 Page:
ISSN:2399-3642
Container-title:Communications Biology
language:en
Short-container-title:Commun Biol

Author:

Clarke Daniel J. B.^ORCID,Marino Giacomo B.^ORCID,Deng Eden Z.,Xie Zhuorui,Evangelista John Erol,Ma’ayan Avi^ORCID

Abstract

AbstractMany biomedical research publications contain gene sets in their supporting tables, and these sets are currently not available for search and reuse. By crawling PubMed Central, the Rummagene server provides access to hundreds of thousands of such mammalian gene sets. So far, we scanned 5,448,589 articles to find 121,237 articles that contain 642,389 gene sets. These sets are served for enrichment analysis, free text, and table title search. Investigating statistical patterns within the Rummagene database, we demonstrate that Rummagene can be used for transcription factor and kinase enrichment analyses, and for gene function predictions. By combining gene set similarity with abstract similarity, Rummagene can find surprising relationships between biological processes, concepts, and named entities. Overall, Rummagene brings to surface the ability to search a massive collection of published biomedical datasets that are currently buried and inaccessible. The Rummagene web application is available at https://rummagene.com.

Funder

U.S. Department of Health & Human Services | NIH | NIH Office of the Director

U.S. Department of Health & Human Services | NIH | National Cancer Institute

U.S. Department of Health & Human Services | NIH | National Institute of Diabetes and Digestive and Kidney Diseases

U.S. Department of Health & Human Services | NIH | NCI | Division of Cancer Epidemiology and Genetics, National Cancer Institute

Publisher

Springer Science and Business Media LLC

Link

https://www.nature.com/articles/s42003-024-06177-7.pdf

Reference52 articles.

1. Manzoni, C. et al. Genome, transcriptome and proteome: the rise of omics data and their integration in biomedical sciences. Brief. Bioinform. 19, 286–302 (2018).

2. Keenan, A. B. et al. ChEA3: transcription factor enrichment analysis by orthogonal omics integration. Nucleic Acids Res. 47, W212–W224 (2019).

3. Lachmann, A. et al. ChEA: transcription factor regulation inferred from integrating genome-wide ChIP-X experiments. Bioinformatics 26, 2438–2444 (2010).

4. Hammal, F., de Langen, P., Bergon, A., Lopez, F. & Ballester, B. ReMap 2022: a database of Human, Mouse, Drosophila and Arabidopsis regulatory regions from an integrative analysis of DNA-binding sequencing experiments. Nucleic Acids Res. 50, D316–D325 (2022).

5. Wilks, C. et al. recount3: summaries and queries for large-scale RNA-seq expression and splicing. Genome Biol. 22, 323 (2021).

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. microRNA-1 Regulates Metabolic Flexibility in Skeletal Muscle via Pyruvate Metabolism;2024-08-10