Reproducible Machine Learning Methods for Lung Cancer Detection Using Computed Tomography Images: Algorithm Development and Validation-Reference-Cited by-同舟云学术

Reproducible Machine Learning Methods for Lung Cancer Detection Using Computed Tomography Images: Algorithm Development and Validation

Published:2020-08-05 Issue:8 Volume:22 Page:e16709
ISSN:1438-8871
Container-title:Journal of Medical Internet Research
language:en
Short-container-title:J Med Internet Res

Author:

Yu Kun-Hsing^ORCID,Lee Tsung-Lu Michael^ORCID,Yen Ming-Hsuan^ORCID,Kou S C^ORCID,Rosen Bruce^ORCID,Chiang Jung-Hsien^ORCID,Kohane Isaac S^ORCID

Abstract

Background Chest computed tomography (CT) is crucial for the detection of lung cancer, and many automated CT evaluation methods have been proposed. Due to the divergent software dependencies of the reported approaches, the developed methods are rarely compared or reproduced. Objective The goal of the research was to generate reproducible machine learning modules for lung cancer detection and compare the approaches and performances of the award-winning algorithms developed in the Kaggle Data Science Bowl. Methods We obtained the source codes of all award-winning solutions of the Kaggle Data Science Bowl Challenge, where participants developed automated CT evaluation methods to detect lung cancer (training set n=1397, public test set n=198, final test set n=506). The performance of the algorithms was evaluated by the log-loss function, and the Spearman correlation coefficient of the performance in the public and final test sets was computed. Results Most solutions implemented distinct image preprocessing, segmentation, and classification modules. Variants of U-Net, VGGNet, and residual net were commonly used in nodule segmentation, and transfer learning was used in most of the classification algorithms. Substantial performance variations in the public and final test sets were observed (Spearman correlation coefficient = .39 among the top 10 teams). To ensure the reproducibility of results, we generated a Docker container for each of the top solutions. Conclusions We compared the award-winning algorithms for lung cancer detection and generated reproducible Docker images for the top solutions. Although convolutional neural networks achieved decent accuracy, there is plenty of room for improvement regarding model generalizability.

Publisher

JMIR Publications Inc.

Subject

Health Informatics

Reference43 articles.

1. Global cancer statistics, 2012

2. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries

3. Reduced Lung-Cancer Mortality with Low-Dose Computed Tomographic Screening

4. Evaluation of Individuals With Pulmonary Nodules: When Is It Lung Cancer?

5. Community Low-Dose CT Lung Cancer Screening: A Prospective Cohort Study

Cited by 40 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Lung Cancer Detection Using Explainable Artificial Intelligence in Medical Diagnosis;Advances in Explainable AI Applications for Smart Cities;2024-01-18

2. Deep Machine Learning for Medical Diagnosis, Application to Lung Cancer Detection: A Review;BioMedInformatics;2024-01-18

3. KFS-Net: Key Features Sampling Network for Lung Nodule Segmentation;Sensing and Imaging;2023-12-14

4. Enhancing Lung Cancer Detection and Classification Using Machine Learning and Deep Learning Techniques: A Comparative Study;2023 International Conference on Networking and Advanced Systems (ICNAS);2023-10-21

5. Early Diagnosis of Lung Nodules With Deep Neural Networks;AI and IoT-Based Technologies for Precision Medicine;2023-10-18