Cohort Selection for Clinical Trials From Longitudinal Patient Records: Text Mining Approach-Reference-Cited by-同舟云学术

Cohort Selection for Clinical Trials From Longitudinal Patient Records: Text Mining Approach

Published:2019-10-31 Issue:4 Volume:7 Page:e15980
ISSN:2291-9694
Container-title:JMIR Medical Informatics
language:en
Short-container-title:JMIR Med Inform

Author:

Spasic Irena^ORCID,Krzeminski Dominik^ORCID,Corcoran Padraig^ORCID,Balinsky Alexander^ORCID

Abstract

Background Clinical trials are an important step in introducing new interventions into clinical practice by generating data on their safety and efficacy. Clinical trials need to ensure that participants are similar so that the findings can be attributed to the interventions studied and not to some other factors. Therefore, each clinical trial defines eligibility criteria, which describe characteristics that must be shared by the participants. Unfortunately, the complexities of eligibility criteria may not allow them to be translated directly into readily executable database queries. Instead, they may require careful analysis of the narrative sections of medical records. Manual screening of medical records is time consuming, thus negatively affecting the timeliness of the recruitment process. Objective Track 1 of the 2018 National Natural Language Processing Clinical Challenge focused on the task of cohort selection for clinical trials, aiming to answer the following question: Can natural language processing be applied to narrative medical records to identify patients who meet eligibility criteria for clinical trials? The task required the participating systems to analyze longitudinal patient records to determine if the corresponding patients met the given eligibility criteria. We aimed to describe a system developed to address this task. Methods Our system consisted of 13 classifiers, one for each eligibility criterion. All classifiers used a bag-of-words document representation model. To prevent the loss of relevant contextual information associated with such representation, a pattern-matching approach was used to extract context-sensitive features. They were embedded back into the text as lexically distinguishable tokens, which were consequently featured in the bag-of-words representation. Supervised machine learning was chosen wherever a sufficient number of both positive and negative instances was available to learn from. A rule-based approach focusing on a small set of relevant features was chosen for the remaining criteria. Results The system was evaluated using microaveraged F measure. Overall, 4 machine algorithms, including support vector machine, logistic regression, naïve Bayesian classifier, and gradient tree boosting (GTB), were evaluated on the training data using 10–fold cross-validation. Overall, GTB demonstrated the most consistent performance. Its performance peaked when oversampling was used to balance the training data. The final evaluation was performed on previously unseen test data. On average, the F measure of 89.04% was comparable to 3 of the top ranked performances in the shared task (91.11%, 90.28%, and 90.21%). With an F measure of 88.14%, we significantly outperformed these systems (81.03%, 78.50%, and 70.81%) in identifying patients with advanced coronary artery disease. Conclusions The holdout evaluation provides evidence that our system was able to identify eligible patients for the given clinical trial with high accuracy. Our approach demonstrates how rule-based knowledge infusion can improve the performance of machine learning algorithms even when trained on a relatively small dataset.

Publisher

JMIR Publications Inc.

Subject

Health Information Management,Health Informatics

Reference54 articles.

1. Clinical trials recruitment planning: A proposed framework from the Clinical Trials Transformation Initiative

2. Unsuccessful trial accrual and human subjects protections: An empirical analysis of recently closed trials

3. Methods to improve recruitment to randomised controlled trials: Cochrane systematic review and meta-analysis

4. Systematic Review and Meta-Analysis of the Magnitude of Structural, Clinical, and Physician and Patient Barriers to Cancer Clinical Trial Participation

Cited by 13 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Clustered Automated Machine Learning (CAML) model for clinical coding multi-label classification;International Journal of Machine Learning and Cybernetics;2024-09-03

2. Word sense disambiguation of acronyms in clinical narratives;Frontiers in Digital Health;2024-02-28

3. The Value of Numbers in Clinical Text Classification;Machine Learning and Knowledge Extraction;2023-07-07

4. Validation and Improvement of a Convolutional Neural Network to Predict the Involved Pathology in a Head and Neck Surgery Cohort;International Journal of Environmental Research and Public Health;2022-09-26

5. Simulation and annotation of global acronyms;Bioinformatics;2022-04-28