An orchestra of machine learning methods reveals landmarks in single-cell data exemplified with aging fibroblasts-Reference-Cited by-同舟云学术

An orchestra of machine learning methods reveals landmarks in single-cell data exemplified with aging fibroblasts

Published:2024-04-17 Issue:4 Volume:19 Page:e0302045
ISSN:1932-6203
Container-title:PLOS ONE
language:en
Short-container-title:PLoS ONE

Author:

Rasbach Lauritz,Caliskan Aylin,Saderi Fatemeh,Dandekar Thomas,Breitenbach Tim^ORCID

Abstract

In this work, a Python framework for characteristic feature extraction is developed and applied to gene expression data of human fibroblasts. Unlabeled feature selection objectively determines groups and minimal gene sets separating groups. ML explainability methods transform the features correlating with phenotypic differences into causal reasoning, supported by further pipeline and visualization tools, allowing user knowledge to boost causal reasoning. The purpose of the framework is to identify characteristic features that are causally related to phenotypic differences of single cells. The pipeline consists of several data science methods enriched with purposeful visualization of the intermediate results in order to check them systematically and infuse the domain knowledge about the investigated process. A specific focus is to extract a small but meaningful set of genes to facilitate causal reasoning for the phenotypic differences. One application could be drug target identification. For this purpose, the framework follows different steps: feature reduction (PFA), low dimensional embedding (UMAP), clustering ((H)DBSCAN), feature correlation (chi-square, mutual information), ML validation and explainability (SHAP, tree explainer). The pipeline is validated by identifying and correctly separating signature genes associated with aging in fibroblasts from single-cell gene expression measurements: PLK3, polo-like protein kinase 3; CCDC88A, Coiled-Coil Domain Containing 88A; STAT3, signal transducer and activator of transcription-3; ZNF7, Zinc Finger Protein 7; SLC24A2, solute carrier family 24 member 2 and lncRNA RP11-372K14.2. The code for the preprocessing step can be found in the GitHub repository https://github.com/AC-PHD/NoLabelPFA, along with the characteristic feature extraction https://github.com/LauritzR/characteristic-feature-extraction.

Funder

Deutsche Forschungsgemeinschaft

Land Bavaria

Publisher

Public Library of Science (PLoS)

Reference69 articles.

1. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.;L McInnes;arXiv,2018

2. Visualizing Data using t-SNE;L van der Maaten;Journal of Machine Learning Research,2008

3. Optimized cell type signatures revealed from single-cell data by combining principal feature analysis, mutual information, and machine learning.;A Caliskan;Computational and Structural Biotechnology Journal.,2023

4. A Unified Approach to Interpreting Model Predictions.;SM Lundberg;arXiv,2017

5. Why should I trust you?" Explaining the predictions of any classifier;MT Ribeiro;Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining,2016.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. DataXflow: Synergizing data-driven modeling with best parameter fit and optimal control – An efficient data analysis for cancer research;Computational and Structural Biotechnology Journal;2024-12