Natural Language Processing and Machine Learning Methods to Characterize Unstructured Patient-Reported Outcomes: Validation Study-Reference-Cited by-同舟云学术

Natural Language Processing and Machine Learning Methods to Characterize Unstructured Patient-Reported Outcomes: Validation Study

Published:2021-11-03 Issue:11 Volume:23 Page:e26777
ISSN:1438-8871
Container-title:Journal of Medical Internet Research
language:en
Short-container-title:J Med Internet Res

Author:

Lu Zhaohua^ORCID,Sim Jin-ah^ORCID,Wang Jade X^ORCID,Forrest Christopher B^ORCID,Krull Kevin R^ORCID,Srivastava Deokumar^ORCID,Hudson Melissa M^ORCID,Robison Leslie L^ORCID,Baker Justin N^ORCID,Huang I-Chan^ORCID

Abstract

Background Assessing patient-reported outcomes (PROs) through interviews or conversations during clinical encounters provides insightful information about survivorship. Objective This study aims to test the validity of natural language processing (NLP) and machine learning (ML) algorithms in identifying different attributes of pain interference and fatigue symptoms experienced by child and adolescent survivors of cancer versus the judgment by PRO content experts as the gold standard to validate NLP/ML algorithms. Methods This cross-sectional study focused on child and adolescent survivors of cancer, aged 8 to 17 years, and caregivers, from whom 391 meaning units in the pain interference domain and 423 in the fatigue domain were generated for analyses. Data were collected from the After Completion of Therapy Clinic at St. Jude Children’s Research Hospital. Experienced pain interference and fatigue symptoms were reported through in-depth interviews. After verbatim transcription, analyzable sentences (ie, meaning units) were semantically labeled by 2 content experts for each attribute (physical, cognitive, social, or unclassified). Two NLP/ML methods were used to extract and validate the semantic features: bidirectional encoder representations from transformers (BERT) and Word2vec plus one of the ML methods, the support vector machine or extreme gradient boosting. Receiver operating characteristic and precision-recall curves were used to evaluate the accuracy and validity of the NLP/ML methods. Results Compared with Word2vec/support vector machine and Word2vec/extreme gradient boosting, BERT demonstrated higher accuracy in both symptom domains, with 0.931 (95% CI 0.905-0.957) and 0.916 (95% CI 0.887-0.941) for problems with cognitive and social attributes on pain interference, respectively, and 0.929 (95% CI 0.903-0.953) and 0.917 (95% CI 0.891-0.943) for problems with cognitive and social attributes on fatigue, respectively. In addition, BERT yielded superior areas under the receiver operating characteristic curve for cognitive attributes on pain interference and fatigue domains (0.923, 95% CI 0.879-0.997; 0.948, 95% CI 0.922-0.979) and superior areas under the precision-recall curve for cognitive attributes on pain interference and fatigue domains (0.818, 95% CI 0.735-0.917; 0.855, 95% CI 0.791-0.930). Conclusions The BERT method performed better than the other methods. As an alternative to using standard PRO surveys, collecting unstructured PROs via interviews or conversations during clinical encounters and applying NLP/ML methods can facilitate PRO assessment in child and adolescent cancer survivors.

Publisher

JMIR Publications Inc.

Subject

Health Informatics

Reference65 articles.

1. Cancer treatment and survivorship statistics, 2019

2. Cancer statistics, 2019

3. Survivors of Childhood Cancer in the United States: Prevalence and Burden of Morbidity

4. Chronic Health Conditions in Adult Survivors of Childhood Cancer

5. Clinical Ascertainment of Health Outcomes Among Adults Treated for Childhood Cancer

Cited by 22 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Insights Into the Patient Experience of Hormone Therapy for Early Breast Cancer Treatment Using Patient Forum Discussions and Natural Language Processing;JCO Clinical Cancer Informatics;2024-08

2. Transforming Hospital Quality Improvement Through Harnessing the Power of Artificial Intelligence;Global Journal on Quality and Safety in Healthcare;2024-08-01

3. Artificial Intelligence and Machine Learning in Cancer Pain: A Systematic Review;Journal of Pain and Symptom Management;2024-08

4. Artificial intelligence to unlock real‐world evidence in clinical oncology: A primer on recent advances;Cancer Medicine;2024-06

5. Using natural language processing in emergency medicine health service research: A systematic review and meta‐analysis;Academic Emergency Medicine;2024-05-16