Small Patient Datasets Reveal Genetic Drivers of Non-Small Cell Lung Cancer Subtypes using a Novel Machine Learning Approach-Reference-Cited by-同舟云学术

Small Patient Datasets Reveal Genetic Drivers of Non-Small Cell Lung Cancer Subtypes using a Novel Machine Learning Approach

Published:2021-07-29 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Cook Moses^ORCID,Qorri Bessi^ORCID,Baskar Amruth,Ziauddin Jalal,Pani Luca^ORCID,Yenkanchi Shashi Bushan,Geraci Joseph^ORCID

Abstract

AbstractThere are many small datasets of significant value in the medical space that are being underutilized. Due to the heterogeneity of complex disorders found in oncology, systems capable of discovering patient subpopulations while elucidating etiologies is of great value as it can indicate leads for innovative drug discovery and development. Here, we report on a machine intelligence-based study that utilized a combination of two small non-small cell lung cancer (NSCLC) datasets consisting of 58 samples of adenocarcinoma (ADC) and squamous cell carcinoma (SCC) and 45 samples from the gene expression analysis of human lung cancer and control samples series (GSE18842). Utilizing a novel machine learning approach, we were able to uncover subpopulations of ADC and SCC while simultaneously extracting which genes, in combination, were significantly involved in defining the subpopulations. An interactive hypothesis-generating interface designed to work with machine learning methods allowed us to explore the hypotheses generated by the unsupervised components of the system. Using these methods, we were able to uncover genes implicated by other methods and accurately discover known subpopulations without being asked, such as different levels of aggressiveness within the SCC and ADC subtypes. Furthermore, PIGX was a novel gene implicated in this study that warrants further study due to its role in breast cancer proliferation. Here we demonstrate the ability to learn from small datasets and reveal well-established properties of NSCLC. These machine learning techniques can reveal the driving factors behind subpopulations of patients altering the approach to drug discovery and development by making precision medicine a reality.

Publisher

Cold Spring Harbor Laboratory

Reference95 articles.

1. Ridge, C.A. , McErlean, A.M. , Ginsberg, M.S. Epidemiology of lung cancer. In Proceedings of Seminars in interventional radiology; p. 93.

2. Refining the treatment of NSCLC according to histological and molecular subtypes;Nature reviews Clinical oncology,2015

3. Discovery and saturation analysis of cancer genes across 21 tumour types

4. Genetic alterations defining NSCLC subtypes and their therapeutic implications

5. Treatment algorithm in 2014 for advanced non-small cell lung cancer: therapy selection by tumour histology and molecular biology;Advances in medical sciences,2014

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An Artificial Intelligence (AI)-Integrated Approach to Enhance Early Detection and Personalized Treatment Strategies in Lung Cancer Among Smokers: A Literature Review;Cureus;2024-08-12

2. Integration of artificial intelligence in lung cancer: Rise of the machine;Cell Reports Medicine;2023-02