Clinical risk prediction using language models: benefits and considerations-Reference-Cited by-同舟云学术

Clinical risk prediction using language models: benefits and considerations

Published:2024-02-27 Issue:9 Volume:31 Page:1856-1864
ISSN:1067-5027
Container-title:Journal of the American Medical Informatics Association
language:en
Short-container-title:

Author:

Acharya Angeela¹,Shrestha Sulabh¹,Chen Anyi²,Conte Joseph²,Avramovic Sanja¹,Sikdar Siddhartha¹,Anastasopoulos Antonios¹,Das Sanmay¹

Affiliation:

1. George Mason University , Fairfax, VA, United States

2. Staten Island Performing Provider System , Staten Island, NY, United States

Abstract

Abstract Objective The use of electronic health records (EHRs) for clinical risk prediction is on the rise. However, in many practical settings, the limited availability of task-specific EHR data can restrict the application of standard machine learning pipelines. In this study, we investigate the potential of leveraging language models (LMs) as a means to incorporate supplementary domain knowledge for improving the performance of various EHR-based risk prediction tasks. Methods We propose two novel LM-based methods, namely “LLaMA2-EHR” and “Sent-e-Med.” Our focus is on utilizing the textual descriptions within structured EHRs to make risk predictions about future diagnoses. We conduct a comprehensive comparison with previous approaches across various data types and sizes. Results Experiments across 6 different methods and 3 separate risk prediction tasks reveal that employing LMs to represent structured EHRs, such as diagnostic histories, results in significant performance improvements when evaluated using standard metrics such as area under the receiver operating characteristic (ROC) curve and precision-recall (PR) curve. Additionally, they offer benefits such as few-shot learning, the ability to handle previously unseen medical concepts, and adaptability to various medical vocabularies. However, it is noteworthy that outcomes may exhibit sensitivity to a specific prompt. Conclusion LMs encompass extensive embedded knowledge, making them valuable for the analysis of EHRs in the context of risk prediction. Nevertheless, it is important to exercise caution in their application, as ongoing safety concerns related to LMs persist and require continuous consideration.

Funder

NSF

Office of Research Computing

George Mason University

Publisher

Oxford University Press (OUP)

Link

https://academic.oup.com/jamia/article-pdf/31/9/1856/58868302/ocae030.pdf

Reference32 articles.

1. Using electronic health records to generate phenotypes for research;Pendergrass;Curr Protoc Hum Genet,2018

2. Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic review;Goldstein;J Am Med Inform Assoc,2017

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Large language models in biomedicine and health: current research landscape and future directions;Journal of the American Medical Informatics Association;2024-08-22

2. Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review;2024-08-12