Improving replicability in single-cell RNA-Seq cell type discovery with Dune-Reference-Cited by-同舟云学术

Improving replicability in single-cell RNA-Seq cell type discovery with Dune

Published:2024-05-24 Issue:1 Volume:25 Page:
ISSN:1471-2105
Container-title:BMC Bioinformatics
language:en
Short-container-title:BMC Bioinformatics

Author:

Roux de Bézieux Hector,Street Kelly,Fischer Stephan,Van den Berge Koen,Chance Rebecca,Risso Davide,Gillis Jesse,Ngai John,Purdom Elizabeth,Dudoit Sandrine

Abstract

Abstract Background Single-cell transcriptome sequencing (scRNA-Seq) has allowed new types of investigations at unprecedented levels of resolution. Among the primary goals of scRNA-Seq is the classification of cells into distinct types. Many approaches build on existing clustering literature to develop tools specific to single-cell. However, almost all of these methods rely on heuristics or user-supplied parameters to control the number of clusters. This affects both the resolution of the clusters within the original dataset as well as their replicability across datasets. While many recommendations exist, in general, there is little assurance that any given set of parameters will represent an optimal choice in the trade-off between cluster resolution and replicability. For instance, another set of parameters may result in more clusters that are also more replicable. Results Here, we propose , a new method for optimizing the trade-off between the resolution of the clusters and their replicability. Our method takes as input a set of clustering results—or partitions—on a single dataset and iteratively merges clusters within each partitions in order to maximize their concordance between partitions. As demonstrated on multiple datasets from different platforms, outperforms existing techniques, that rely on hierarchical merging for reducing the number of clusters, in terms of replicability of the resultant merged clusters as well as concordance with ground truth. is available as an R package on Bioconductor: https://www.bioconductor.org/packages/release/bioc/html/Dune.html. Conclusions Cluster refinement by helps improve the robustness of any clustering analysis and reduces the reliance on tuning parameters. This method provides an objective approach for borrowing information across multiple clusterings to generate replicable clusters most likely to represent common biological features across multiple datasets.

Funder

Fonds Wetenschappelijk Onderzoek

National Institutes of Health

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1186/s12859-024-05814-6.pdf

Reference40 articles.

1. Svensson V, da Veiga Beltrame E. A curated database reveals trends in single cell transcriptomics. bioRxiv; 2019. pp. 742304. https://doi.org/10.1101/742304.

2. Kiselev VY, Kirschner K, Schaub MT, Andrews T, Yiu A, Chandra T, Natarajan KN, Reik W, Barahona M, Green AR, Hemberg M. SC3: consensus clustering of single-cell RNA-seq data. Nat Methods. 2017;14(5):483–6. https://doi.org/10.1038/nmeth.4236.

3. Stuart T, Butler A, Hoffman P, Hafemeister C, Papalexi E, Mauck WM, Hao Y, Stoeckius M, Smibert P, Satija R. Comprehensive integration of single-cell data. Cell. 2019;177(7):1888. https://doi.org/10.1016/j.cell.2019.05.031.

4. Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L, Steemers FJ, Trapnell C, Shendure J. The single-cell transcriptional landscape of mammalian organogenesis. Nature. 2019;566(7745):496–502. https://doi.org/10.1038/s41586-019-0969-x.

5. Duò A, Robinson MD, Soneson C. A systematic performance evaluation of clustering methods for single-cell RNA-seq data. F1000Research. 2018;7:377–82. https://doi.org/10.5256/f1000research.17093.r36544.