Identifying Information Gaps in Electronic Health Records by Using Natural Language Processing: Gynecologic Surgery History Identification-Reference-Cited by-同舟云学术

Identifying Information Gaps in Electronic Health Records by Using Natural Language Processing: Gynecologic Surgery History Identification

Published:2022-01-28 Issue:1 Volume:24 Page:e29015
ISSN:1438-8871
Container-title:Journal of Medical Internet Research
language:en
Short-container-title:J Med Internet Res

Author:

Moon Sungrim^ORCID,Carlson Luke A^ORCID,Moser Ethan D^ORCID,Agnikula Kshatriya Bhavani Singh^ORCID,Smith Carin Y^ORCID,Rocca Walter A^ORCID,Gazzuola Rocca Liliana^ORCID,Bielinski Suzette J^ORCID,Liu Hongfang^ORCID,Larson Nicholas B^ORCID

Abstract

Background Electronic health records (EHRs) are a rich source of longitudinal patient data. However, missing information due to clinical care that predated the implementation of EHR system(s) or care that occurred at different medical institutions impedes complete ascertainment of a patient’s medical history. Objective This study aimed to investigate information discrepancies and to quantify information gaps by comparing the gynecological surgical history extracted from an EHR of a single institution by using natural language processing (NLP) techniques with the manually curated surgical history information through chart review of records from multiple independent regional health care institutions. Methods To facilitate high-throughput evaluation, we developed a rule-based NLP algorithm to detect gynecological surgery history from the unstructured narrative of the Mayo Clinic EHR. These results were compared to a gold standard cohort of 3870 women with gynecological surgery status adjudicated using the Rochester Epidemiology Project medical records–linkage system. We quantified and characterized the information gaps observed that led to misclassification of the surgical status. Results The NLP algorithm achieved precision of 0.85, recall of 0.82, and F1-score of 0.83 in the test set (n=265) relative to outcomes abstracted from the Mayo EHR. This performance attenuated when directly compared to the gold standard (precision 0.79, recall 0.76, and F1-score 0.76), with the majority of misclassifications being false negatives in nature. We then applied the algorithm to the remaining patients (n=3340) and identified 2 types of information gaps through error analysis. First, 6% (199/3340) of women in this study had no recorded surgery information or partial information in the EHR. Second, 4.3% (144/3340) of women had inconsistent or inaccurate information within the clinical narrative owing to misinterpreted information, erroneous “copy and paste,” or incorrect information provided by patients. Additionally, the NLP algorithm misclassified the surgery status of 3.6% (121/3340) of women. Conclusions Although NLP techniques were able to adequately recreate the gynecologic surgical status from the clinical narrative, missing or inaccurately reported and recorded information resulted in much of the misclassification observed. Therefore, alternative approaches to collect or curate surgical history are needed.

Publisher

JMIR Publications Inc.

Subject

Health Informatics

Reference25 articles.

1. Ovarian Conservation at the Time of Hysterectomy and Long-Term Health Outcomes in the Nurses’ Health Study

2. Long-Term Mortality Associated With Oophorectomy Compared With Ovarian Conservation in the Nurses' Health Study

3. Survival patterns after oophorectomy in premenopausal women: a population-based cohort study

4. Fracture risk following bilateral oophorectomy

5. Accelerated Accumulation of Multimorbidity After Bilateral Oophorectomy: A Population-Based Cohort Study

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Impact of Stressful Life Events on Preventive Colon Cancer Screening Adherence: Large Language Model Approach (Preprint);2024-09-06

2. Blockchains in health information systems: A literature review on use cases and status of implementation of blockchains for electronic health records;Human Systems Management;2024-09-05

3. Evaluation of ChatGPT for Pelvic Floor Surgery Counseling;Urogynecology;2024-03

4. Incorporating reproductive system history data into cardiovascular nursing research to advance women’s health;European Journal of Cardiovascular Nursing;2024-01-10

5. A Blockchain Patient-Centric Records Framework for Older Adult Healthcare;Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering;2023-12-15