Labelling chest x-ray reports using an open-source NLP and ML tool for text data binary classification-Reference-Cited by-同舟云学术

Labelling chest x-ray reports using an open-source NLP and ML tool for text data binary classification

Published:2019-11-22 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Towfighi Sohrab,Agarwal Arnav,Mak Denise Y. F.,Verma Amol

Abstract

AbstractThe chest x-ray is a commonly requested diagnostic test on internal medicine wards which can diagnose many acute pathologies needing intervention. We developed a natural language processing (NLP) and machine learning (ML) model to identify the presence of opacities or endotracheal intubation on chest x-rays using only the radiology report. This a preliminary report of our work and findings. Using the General Medicine Inpatient Initiative (GEMINI) dataset, housing inpatient clinical and administrative data from 7 major hospitals, we retrieved 1000 plain film radiology reports which were classified according to 4 labels by an internal medicine resident. NLP/ML models were developed to identify the following on the radiograph reports: the report is that of a chest x-ray, there is definite absence of an opacity, there is definite presence of an opacity, the report is a follow-up report with minimal details in its text, and there is an endotracheal tube in place. Our NLP/ML model development methodology included a random search of either TF-IDF or bag-of-words for vectorization along with random search of various ML models. Our Python programming scripts were made publicly available on GitHub to allow other parties to train models using their own text data. 100 randomly generated ML pipelines were compared using 10-fold cross validation on 75% of the data, while 25% of the data was left out for generalizability testing. With respect to the question of whether a chest x-ray definitely lacks an opacity, the model’s performance metrics were accuracy of 0.84, precision of 0.94, recall of 0.81, and receiver operating characteristic area under curve of 0.86. Model performance was worse when trained against a highly imbalanced dataset despite the use of an advanced oversampling technique.

Publisher

Cold Spring Harbor Laboratory

Reference20 articles.

1. Clinical utility of chest roentgenograms

2. A Retrospective Analysis of the Clinical Impact of 939 Chest Radiographs Using the Medical Records

3. Patient characteristics, resource use and outcomes associated with general internal medicine hospital care: the General Medicine Inpatient Initiative (GEMINI) retrospective cohort study

4. Prevalence and Costs of Discharge Diagnoses in Inpatient General Internal Medicine: a Multi-center Cross-sectional Study

5. Mortality, morbidity, and disease severity of patients with aspiration pneumonia

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Recent Trends in Deep Learning for Conversational AI;Artificial Intelligence;2023-11-03

2. A Web-Based Platform for the Automatic Stratification of ARDS Severity;Diagnostics;2023-03-01

3. A method to identify p62’s UBA domain interacting proteins;Biological Procedures Online;2003-02