Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow-Reference-Cited by-同舟云学术

Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow

Published:2023-02-26 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Rao Arya^ORCID,Pang Michael,Kim John,Kamineni Meghana,Lie Winston,Prasad Anoop K.,Landman Adam,Dreyer Keith J,Succi Marc D.^ORCID

Abstract

AbstractIMPORTANCELarge language model (LLM) artificial intelligence (AI) chatbots direct the power of large training datasets towards successive, related tasks, as opposed to single-ask tasks, for which AI already achieves impressive performance. The capacity of LLMs to assist in the full scope of iterative clinical reasoning via successive prompting, in effect acting as virtual physicians, has not yet been evaluated.OBJECTIVETo evaluate ChatGPT’s capacity for ongoing clinical decision support via its performance on standardized clinical vignettes.DESIGNWe inputted all 36 published clinical vignettes from the Merck Sharpe & Dohme (MSD) Clinical Manual into ChatGPT and compared accuracy on differential diagnoses, diagnostic testing, final diagnosis, and management based on patient age, gender, and case acuity.SETTINGChatGPT, a publicly available LLMPARTICIPANTSClinical vignettes featured hypothetical patients with a variety of age and gender identities, and a range of Emergency Severity Indices (ESIs) based on initial clinical presentation.EXPOSURESMSD Clinical Manual vignettesMAIN OUTCOMES AND MEASURESWe measured the proportion of correct responses to the questions posed within the clinical vignettes tested.RESULTSChatGPT achieved 71.7% (95% CI, 69.3% to 74.1%) accuracy overall across all 36 clinical vignettes. The LLM demonstrated the highest performance in making a final diagnosis with an accuracy of 76.9% (95% CI, 67.8% to 86.1%), and the lowest performance in generating an initial differential diagnosis with an accuracy of 60.3% (95% CI, 54.2% to 66.6%). Compared to answering questions about general medical knowledge, ChatGPT demonstrated inferior performance on differential diagnosis (β=-15.8%, p<0.001) and clinical management (β=-7.4%, p=0.02) type questions.CONCLUSIONS AND RELEVANCEChatGPT achieves impressive accuracy in clinical decision making, with particular strengths emerging as it has more clinical information at its disposal.

Publisher

Cold Spring Harbor Laboratory

Reference27 articles.

1. Artificial intelligence in healthcare

2. Chatbot for Health Care and Oncology Applications Using Artificial Intelligence and Machine Learning: Systematic Review

3. RadTranslate: An Artificial Intelligence–Powered Intervention for Urgent Imaging to Enhance Care Equity for Patients With Limited English Proficiency During the COVID-19 Pandemic

4. Prediction of oxygen requirement in patients with COVID-19 using a pre-trained chest radiograph xAI model: efficient development of auditable risk prediction models via a fine-tuning approach

5. Multi-population generalizability of a deep learning-based chest radiograph severity score for COVID-19

Cited by 86 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Programming Chatbots Using Natural Language: Generating Cervical Spine MRI Impressions;Cureus;2024-09-14

2. “Hospice Care Could Be a Compassionate Choice”: ChatGPT Responses to Questions About Decision Making in Advanced Cancer;Journal of Palliative Medicine;2024-09-12

3. Assessing Artificial Intelligence–Generated Responses to Urology Patient In-Basket Messages;Urology Practice;2024-09

4. Predicting the risk category of thymoma with machine learning-based computed tomography radiomics signatures and their between-imaging phase differences;Scientific Reports;2024-08-19

5. Large language models (LLMs): survey, technical frameworks, and future challenges;Artificial Intelligence Review;2024-08-18