Real-World Performance of Large Language Models in Emergency Department Chest Pain Triage

Author:

Meng Xiangbin,Ji Jia-ming,Yan Xiangyu,Xu Hua,gao Jun,Wang Junhong,Wang Jingjia,Wang Xuliang,Wang Yuan-geng-shuo,Wang Wenyao,Chen Jing,Zhang Kuo,Liu Da,Qiu Zifeng,Li Muzi,Shao Chunli,Yang Yaodong,Tang Yi-Da

Abstract

AbstractBackgroundLarge Language Models (LLMs) are increasingly being explored for medical applications, particularly in emergency triage where rapid and accurate decision-making is crucial. This study evaluates the diagnostic performance of two prominent Chinese LLMs, “Tongyi Qianwen” and “Lingyi Zhihui,” alongside a newly developed model, MediGuide-14B, comparing their effectiveness with human medical experts in emergency chest pain triage.MethodsConducted at Peking University Third Hospital’s emergency centers from June 2021 to May 2023, this retrospective study involved 11,428 patients with chest pain symptoms. Data were extracted from electronic medical records, excluding diagnostic test results, and used to assess the models and human experts in a double-blind setup. The models’ performances were evaluated based on their accuracy, sensitivity, and specificity in diagnosing Acute Coronary Syndrome (ACS).Findings“Lingyi Zhihui” demonstrated a diagnostic accuracy of 76.40%, sensitivity of 90.99%, and specificity of 70.15%. “Tongyi Qianwen” showed an accuracy of 61.11%, sensitivity of 91.67%, and specificity of 47.95%. MediGuide-14B outperformed these models with an accuracy of 84.52%, showcasing high sensitivity and commendable specificity. Human experts achieved higher accuracy (86.37%) and specificity (89.26%) but lower sensitivity compared to the LLMs. The study also highlighted the potential of LLMs to provide rapid triage decisions, significantly faster than human experts, though with varying degrees of reliability and completeness in their recommendations.InterpretationThe study confirms the potential of LLMs in enhancing emergency medical diagnostics, particularly in settings with limited resources. MediGuide-14B, with its tailored training for medical applications, demonstrates considerable promise for clinical integration. However, the variability in performance underscores the need for further fine-tuning and contextual adaptation to improve reliability and efficacy in medical applications. Future research should focus on optimizing LLMs for specific medical tasks and integrating them with conventional medical systems to leverage their full potential in real-world settings.

Publisher

Cold Spring Harbor Laboratory

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3