Scaling cross-tissue single-cell annotation models-Reference-Cited by-同舟云学术

Scaling cross-tissue single-cell annotation models

Published:2023-10-10 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Fischer Felix,Fischer David S.^ORCID,Biederstedt Evan,Villani Alexandra-Chloé,Theis Fabian J.^ORCID

Abstract

Identifying cellular identities (both novel and well-studied) is one of the key use cases in single-cell transcriptomics. While supervised machine learning has been leveraged to automate cell annotation predictions for some time, there has been relatively little progress both in scaling neural networks to large data sets and in constructing models that generalize well across diverse tissues and biological contexts up to whole organisms. Here, we propose scTab, an automated, feature-attention-based cell type prediction model specific to tabular data, and train it using a novel data augmentation scheme across a large corpus of single-cell RNA-seq observations (22.2 million human cells in total). In addition, scTab leverages deep ensembles for uncertainty quantification. Moreover, we account for ontological relationships between labels in the model evaluation to accommodate for differences in annotation granularity across datasets. On this large-scale corpus, we show that cross-tissue annotation requires nonlinear models and that the performance of scTab scales in terms of training dataset size as well as model size - demonstrating the advantage of scTab over current state-of-the-art linear models in this context. Additionally, we show that the proposed data augmentation schema improves model generalization. In summary, we introduce a de novo cell type prediction model for single-cell RNA-seq data that can be trained across a large-scale collection of curated datasets from a diverse selection of human tissues and demonstrate the benefits of using deep learning methods in this paradigm. Our codebase, training data, and model checkpoints are publicly available athttps://github.com/theislab/scTabto further enable rigorous benchmarks of foundation models for single-cell RNA-seq data.

Publisher

Cold Spring Harbor Laboratory

Reference43 articles.

1. Benchmarking atlas-level data integration in single-cell genomics

2. Current best practices in single‐cell RNA‐seq analysis: a tutorial

3. Best practices for single-cell analysis across modalities;Nat. Rev. Genet,2023

4. Orchestrating single-cell analysis with Bioconductor;Nat. Methods,2020

5. Abdelaal, T. et al. A comparison of automatic cell identification methods for single-cell RNA sequencing data. Genome Biol. 20, 194 (2019).

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Transformers in single-cell omics: a review and new perspectives;Nature Methods;2024-08

2. scBOL: a universal cell type identification framework for single-cell and spatial transcriptomics data;Briefings in Bioinformatics;2024-03-27

3. Parameter-Efficient Fine-Tuning Enhances Adaptation of Single Cell Large Language Model for Cell Type Identification;2024-01-30

4. CZ CELL×GENE Discover: A single-cell data platform for scalable exploration, analysis and modeling of aggregated data;2023-11-02