Automated NLP Extraction of Clinical Rationale for Treatment Discontinuation in Breast Cancer-Reference-Cited by-同舟云学术

Automated NLP Extraction of Clinical Rationale for Treatment Discontinuation in Breast Cancer

Published:2021-05 Issue:5 Volume: Page:550-560
ISSN:2473-4276
Container-title:JCO Clinical Cancer Informatics
language:en
Short-container-title:JCO Clinical Cancer Informatics

Author:

Alkaitis Matthew S.¹²^ORCID,Agrawal Monica N.¹^ORCID,Riely Gregory J.³⁴^ORCID,Razavi Pedram³⁴,Sontag David¹

Affiliation:

1. CSAIL & IMES, Massachusetts Institute of Technology, Cambridge, MA

2. Harvard Medical School, Boston, MA

3. Memorial Sloan Kettering Cancer Center, New York, NY

4. Weill-Cornell Medical College, New York, NY

Abstract

PURPOSE Key oncology end points are not routinely encoded into electronic medical records (EMRs). We assessed whether natural language processing (NLP) can abstract treatment discontinuation rationale from unstructured EMR notes to estimate toxicity incidence and progression-free survival (PFS). METHODS We constructed a retrospective cohort of 6,115 patients with early-stage and 701 patients with metastatic breast cancer initiating care at Memorial Sloan Kettering Cancer Center from 2008 to 2019. Each cohort was divided into training (70%), validation (15%), and test (15%) subsets. Human abstractors identified the clinical rationale associated with treatment discontinuation events. Concatenated EMR notes were used to train high-dimensional logistic regression and convolutional neural network models. Kaplan-Meier analyses were used to compare toxicity incidence and PFS estimated by our NLP models to estimates generated by manual labeling and time-to-treatment discontinuation (TTD). RESULTS Our best high-dimensional logistic regression models identified toxicity events in early-stage patients with an area under the curve of the receiver-operator characteristic of 0.857 ± 0.014 (standard deviation) and progression events in metastatic patients with an area under the curve of 0.752 ± 0.027 (standard deviation). NLP-extracted toxicity incidence and PFS curves were not significantly different from manually extracted curves ( P = .95 and P = .67, respectively). By contrast, TTD overestimated toxicity in early-stage patients ( P < .001) and underestimated PFS in metastatic patients ( P < .001). Additionally, we tested an extrapolation approach in which 20% of the metastatic cohort were labeled manually, and NLP algorithms were used to abstract the remaining 80%. This extrapolated outcomes approach resolved PFS differences between receptor subtypes ( P < .001 for hormone receptor+/human epidermal growth factor receptor 2− v human epidermal growth factor receptor 2+ v triple-negative) that could not be resolved with TTD. CONCLUSION NLP models are capable of abstracting treatment discontinuation rationale with minimal manual labeling.

Publisher

American Society of Clinical Oncology (ASCO)

Subject

General Medicine

Link

https://ascopubs.org/doi/pdfdirect/10.1200/CCI.20.00139

Reference58 articles.

1. Impact of the HITECH financial incentives on EHR adoption in small, physician-owned practices

2. Validation of time to treatment change (TTC) as a surrogate end-point of progression free survival (PFS) for observational trials in metastatic breast cancer patients (MBC): The GIM-13 AMBRA study.

3. Analysis of time-to-treatment discontinuation of targeted therapy, immunotherapy, and chemotherapy in clinical trials of patients with non-small-cell lung cancer

4. New response evaluation criteria in solid tumours: Revised RECIST guideline (version 1.1)

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Artificial intelligence auxiliary diagnosis and treatment system for breast cancer in developing countries;Journal of X-Ray Science and Technology;2024-01-05

2. The emergent role of artificial intelligence, natural learning processing, and large language models in higher education and research;Research in Social and Administrative Pharmacy;2023-08

3. Natural Language Processing for Breast Imaging: A Systematic Review;Diagnostics;2023-04-14

4. Exploration of biomedical knowledge for recurrent glioblastoma using natural language processing deep learning models;BMC Medical Informatics and Decision Making;2022-10-13