Assessing the research landscape and clinical utility of large language models: a scoping review-Reference-Cited by-同舟云学术

Assessing the research landscape and clinical utility of large language models: a scoping review

Published:2024-03-12 Issue:1 Volume:24 Page:
ISSN:1472-6947
Container-title:BMC Medical Informatics and Decision Making
language:en
Short-container-title:BMC Med Inform Decis Mak

Author:

Park Ye-Jean,Pillai Abhinav,Deng Jiawen,Guo Eddie,Gupta Mehul,Paget Mike,Naugler Christopher

Abstract

Abstract Importance Large language models (LLMs) like OpenAI’s ChatGPT are powerful generative systems that rapidly synthesize natural language responses. Research on LLMs has revealed their potential and pitfalls, especially in clinical settings. However, the evolving landscape of LLM research in medicine has left several gaps regarding their evaluation, application, and evidence base. Objective This scoping review aims to (1) summarize current research evidence on the accuracy and efficacy of LLMs in medical applications, (2) discuss the ethical, legal, logistical, and socioeconomic implications of LLM use in clinical settings, (3) explore barriers and facilitators to LLM implementation in healthcare, (4) propose a standardized evaluation framework for assessing LLMs’ clinical utility, and (5) identify evidence gaps and propose future research directions for LLMs in clinical applications. Evidence review We screened 4,036 records from MEDLINE, EMBASE, CINAHL, medRxiv, bioRxiv, and arXiv from January 2023 (inception of the search) to June 26, 2023 for English-language papers and analyzed findings from 55 worldwide studies. Quality of evidence was reported based on the Oxford Centre for Evidence-based Medicine recommendations. Findings Our results demonstrate that LLMs show promise in compiling patient notes, assisting patients in navigating the healthcare system, and to some extent, supporting clinical decision-making when combined with human oversight. However, their utilization is limited by biases in training data that may harm patients, the generation of inaccurate but convincing information, and ethical, legal, socioeconomic, and privacy concerns. We also identified a lack of standardized methods for evaluating LLMs’ effectiveness and feasibility. Conclusions and relevance This review thus highlights potential future directions and questions to address these limitations and to further explore LLMs’ potential in enhancing healthcare delivery.

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1186/s12911-024-02459-6.pdf

Reference81 articles.

1. Yang X, Chen A, PourNejatian N, Shin HC, Smith KE, Parisien C, et al. A large language model for electronic health records. NPJ Digit Med. 2022;5(1):194.

2. OpenAI. Introducing ChatGPT [Internet]. [cited 2023 May 2]. Available from: https://openai.com/blog/chatgpt.

3. Devlin J, Chang MW, Lee K, Toutanova K, BERT. Pre-training of deep bidirectional Transformers for language understanding [Internet]. arXiv. 2018. Available from: https://arxiv.org/abs/1810.04805.

4. Levine DM, Tuwani R, Kompa B, Varma A, Finlayson SG, Mehrotra A et al. The Diagnostic and Triage Accuracy of the GPT-3 Artificial Intelligence Model [Internet]. medRxiv. 2023. https://doi.org/10.1101/2023.01.30.23285067.

5. Stewart J, Lu J, Goudie A, Arendts G, Meka SA, Freeman S et al. Applications of natural language processing at emergency department triage: A systematic review [Internet]. bioRxiv. 2022. https://doi.org/10.1101/2022.12.20.22283735.

Cited by 18 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Visual-Textual Integration in LLMs for Medical Diagnosis: A Quantitative Analysis;2024-09-03

2. Integrating machine learning and artificial intelligence in life-course epidemiology: pathways to innovative public health solutions;BMC Medicine;2024-09-02

3. Unlocking the potential of advanced large language models in medication review and reconciliation: A proof-of-concept investigation;Exploratory Research in Clinical and Social Pharmacy;2024-09

4. Large language models and artificial intelligence chatbots in vascular surgery;Seminars in Vascular Surgery;2024-09

5. How to critically appraise and direct the trajectory of AI development and application in oncology;ESMO Real World Data and Digital Oncology;2024-09