How does ChatGPT4 preform on Non-English National Medical Licensing Examination? An Evaluation in Chinese Language (Preprint)-Reference-Cited by-同舟云学术

How does ChatGPT4 preform on Non-English National Medical Licensing Examination? An Evaluation in Chinese Language (Preprint)

Published:2023-05-03 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Fang Changchang,Ling Jitao,Zhou Jing,Wang Yue,Liu Xiaolin,Jiang Yuan,Wu Yifan,Chen Yixuan,Zhu Zhichen,Ma Jianyong,Yan Ziwei,Yu Peng,Liu Xiao

Abstract

BACKGROUND

ChatGPT, an artificial intelligence (AI) system powered by large-scale language models, has garnered significant interest in the healthcare. Its performance dependent on the quality and amount of training data available for specific language. This study aims to assess the of ChatGPT's ability in medical education and clinical decision-making within the Chinese context.

OBJECTIVE

Evaluate ChatGPT's performance on the Chinese NMLE conducted within the Chinese context.

METHODS

We utilized a dataset from the Chinese National Medical Licensing Examination (NMLE) to assess ChatGPT-4's proficiency in medical knowledge within the Chinese language. Performance indicators, including score, accuracy, and concordance (confirmation of answers through explanation), were employed to evaluate ChatGPT's effectiveness in both original and encoded medical questions. Additionally, we translated the original Chinese questions into English to explore potential avenues for improvement.

RESULTS

ChatGPT scored 442/600 for original questions in Chinese, surpassing the passing threshold of 360/600. However, ChatGPT demonstrated reduced accuracy in addressing open-ended questions, with an overall accuracy rate of 47.7%. Despite this, ChatGPT displayed commendable consistency, achieving a 75% concordance rate across all case analysis questions. Moreover, translating Chinese case analysis questions into English yielded only marginal improvements in ChatGPT's performance (P =0.728).

CONCLUSIONS

ChatGPT exhibits remarkable precision and reliability when handling the NMLE in Chinese language. Translation of NMLE questions from Chinese to English does not yield an improvement in ChatGPT's performance.

Publisher

JMIR Publications Inc.

Reference16 articles.

1. Current status and applications of Artificial Intelligence (AI) in medical field: An overview

2. Artificial Intelligence (AI) applications in orthopaedics: An innovative technology to embrace

3. Information and Artificial Intelligence

4. Some ethical and legal consequences of the application of artificial intelligence in the field of medicine

5. The Inevitable Application of Big Data to Health Care

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Capability of GPT-4V(ision) in the Japanese National Medical Licensing Examination: Evaluation Study (Preprint);2023-11-08