Author:
Adamson Blythe,Waskom Michael,Blarre Auriane,Kelly Jonathan,Krismer Konstantin,Nemeth Sheila,Gippetti James,Ritten John,Harrison Katherine,Ho George,Linzmayer Robin,Bansal Tarun,Wilkinson Samuel,Amster Guy,Estola Evan,Benedum Corey M.,Fidyk Erin,Estévez Melissa,Shapiro Will,Cohen Aaron B.
Abstract
Background: As artificial intelligence (AI) continues to advance with breakthroughs in natural language processing (NLP) and machine learning (ML), such as the development of models like OpenAI’s ChatGPT, new opportunities are emerging for efficient curation of electronic health records (EHR) into real-world data (RWD) for evidence generation in oncology. Our objective is to describe the research and development of industry methods to promote transparency and explainability.Methods: We applied NLP with ML techniques to train, validate, and test the extraction of information from unstructured documents (e.g., clinician notes, radiology reports, lab reports, etc.) to output a set of structured variables required for RWD analysis. This research used a nationwide electronic health record (EHR)-derived database. Models were selected based on performance. Variables curated with an approach using ML extraction are those where the value is determined solely based on an ML model (i.e. not confirmed by abstraction), which identifies key information from visit notes and documents. These models do not predict future events or infer missing information.Results: We developed an approach using NLP and ML for extraction of clinically meaningful information from unstructured EHR documents and found high performance of output variables compared with variables curated by manually abstracted data. These extraction methods resulted in research-ready variables including initial cancer diagnosis with date, advanced/metastatic diagnosis with date, disease stage, histology, smoking status, surgery status with date, biomarker test results with dates, and oral treatments with dates.Conclusion: NLP and ML enable the extraction of retrospective clinical data in EHR with speed and scalability to help researchers learn from the experience of every person with cancer.
Subject
Pharmacology (medical),Pharmacology
Reference51 articles.
1. What's in a summary? Laying the groundwork for advances in hospital-course summarization;Adams;Proc. Conf.,2021
2. Cancer immunotherapy use and effectiveness in real-world patients living with HIV;Adamson,2019
3. Tifti: A framework for extracting drug intervals from longitudinal clinic notes;Agrawal,2018
4. PPM8 A machine learning model for cancer biomarker identification in electronic health records;Ambwani;Value Health,2019
Cited by
3 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献