How does ChatGPT-4 preform on non-English national medical licensing examination? An evaluation in Chinese language-Reference-Cited by-同舟云学术

How does ChatGPT-4 preform on non-English national medical licensing examination? An evaluation in Chinese language

Published:2023-12-01 Issue:12 Volume:2 Page:e0000397
ISSN:2767-3170
Container-title:PLOS Digital Health
language:en
Short-container-title:PLOS Digit Health

Author:

Fang Changchang,Wu Yuting,Fu Wanying,Ling Jitao,Wang Yue,Liu Xiaolin,Jiang Yuan,Wu Yifan,Chen Yixuan,Zhou Jing,Zhu Zhichen,Yan Zhiwei,Yu Peng,Liu Xiao^ORCID

Abstract

ChatGPT, an artificial intelligence (AI) system powered by large-scale language models, has garnered significant interest in healthcare. Its performance dependent on the quality and quantity of training data available for a specific language, with the majority of it being in English. Therefore, its effectiveness in processing the Chinese language, which has fewer data available, warrants further investigation. This study aims to assess the of ChatGPT’s ability in medical education and clinical decision-making within the Chinese context. We utilized a dataset from the Chinese National Medical Licensing Examination (NMLE) to assess ChatGPT-4’s proficiency in medical knowledge in Chinese. Performance indicators, including score, accuracy, and concordance (confirmation of answers through explanation), were employed to evaluate ChatGPT’s effectiveness in both original and encoded medical questions. Additionally, we translated the original Chinese questions into English to explore potential avenues for improvement. ChatGPT scored 442/600 for original questions in Chinese, surpassing the passing threshold of 360/600. However, ChatGPT demonstrated reduced accuracy in addressing open-ended questions, with an overall accuracy rate of 47.7%. Despite this, ChatGPT displayed commendable consistency, achieving a 75% concordance rate across all case analysis questions. Moreover, translating Chinese case analysis questions into English yielded only marginal improvements in ChatGPT’s performance (p = 0.728). ChatGPT exhibits remarkable precision and reliability when handling the NMLE in Chinese. Translation of NMLE questions from Chinese to English does not yield an improvement in ChatGPT’s performance.

Publisher

Public Library of Science (PLoS)

Reference15 articles.

1. Current status and applications of Artificial Intelligence (AI) in medical field: An overview;A Haleem;Current Medicine Research and Practice,2019

2. Artificial Intelligence (AI) applications in orthopaedics: An innovative technology to embrace;A Haleem;Journal of Clinical Orthopaedics and Trauma,2019

3. Information and artificial intelligence;S Jha;Journal of the American College of Radiology,2018

4. The inevitable application of big data to health care;TB Murdoch;JAMA,2013

Cited by 14 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ChatGPT as a global doctor: a rapid review of its performance on national licensing medical examination (Preprint);2024-08-29

2. Large language models in healthcare: from a systematic review on medical examinations to a comparative analysis on fundamentals of robotic surgery online test;Artificial Intelligence Review;2024-08-06

3. Will ChatGPT be Useful for Korean Neurologists in Clinical Practice?;Journal of the Korean Neurological Association;2024-08-01

4. Evaluating the competency of ChatGPT in MRCP Part 1 and a systematic literature review of its capabilities in postgraduate medical assessments;PLOS ONE;2024-07-31

5. Performance of ChatGPT Across Different Versions in Medical Licensing Examinations Worldwide: Systematic Review and Meta-Analysis;Journal of Medical Internet Research;2024-07-25