Performance of ChatGPT on Chinese national medical licensing examinations: a five-year examination evaluation study for physicians, pharmacists and nurses-Reference-Cited by-同舟云学术

Performance of ChatGPT on Chinese national medical licensing examinations: a five-year examination evaluation study for physicians, pharmacists and nurses

Published:2024-02-14 Issue:1 Volume:24 Page:
ISSN:1472-6920
Container-title:BMC Medical Education
language:en
Short-container-title:BMC Med Educ

Author:

Zong Hui^ORCID,Li Jiakun^ORCID,Wu Erman^ORCID,Wu Rongrong^ORCID,Lu Junyu,Shen Bairong^ORCID

Abstract

Abstract Background Large language models like ChatGPT have revolutionized the field of natural language processing with their capability to comprehend and generate textual content, showing great potential to play a role in medical education. This study aimed to quantitatively evaluate and comprehensively analysis the performance of ChatGPT on three types of national medical examinations in China, including National Medical Licensing Examination (NMLE), National Pharmacist Licensing Examination (NPLE), and National Nurse Licensing Examination (NNLE). Methods We collected questions from Chinese NMLE, NPLE and NNLE from year 2017 to 2021. In NMLE and NPLE, each exam consists of 4 units, while in NNLE, each exam consists of 2 units. The questions with figures, tables or chemical structure were manually identified and excluded by clinician. We applied direct instruction strategy via multiple prompts to force ChatGPT to generate the clear answer with the capability to distinguish between single-choice and multiple-choice questions. Results ChatGPT failed to pass the accuracy threshold of 0.6 in any of the three types of examinations over the five years. Specifically, in the NMLE, the highest recorded accuracy was 0.5467, which was attained in both 2018 and 2021. In the NPLE, the highest accuracy was 0.5599 in 2017. In the NNLE, the most impressive result was shown in 2017, with an accuracy of 0.5897, which is also the highest accuracy in our entire evaluation. ChatGPT’s performance showed no significant difference in different units, but significant difference in different question types. ChatGPT performed well in a range of subject areas, including clinical epidemiology, human parasitology, and dermatology, as well as in various medical topics such as molecules, health management and prevention, diagnosis and screening. Conclusions These results indicate ChatGPT failed the NMLE, NPLE and NNLE in China, spanning from year 2017 to 2021. but show great potential of large language models in medical education. In the future high-quality medical data will be required to improve the performance.

Funder

National Natural Science Foundation of China

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1186/s12909-024-05125-7.pdf

Reference29 articles.

1. Bhinder B, et al. Artificial Intelligence in Cancer Research and Precision Medicine. Cancer Discov. 2021;11(4):900–15.

2. Moor M, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616(7956):259–65.

3. van Dis EAM, et al. ChatGPT: five priorities for research. Nature. 2023;614(7947):224–6.

4. Sarink MJ et al. A study on the performance of ChatGPT in infectious diseases clinical consultation. Clin Microbiol Infect, 2023.

5. Lee TC et al. ChatGPT Answers Common Patient Questions About Colonoscopy. Gastroenterology, 2023.

Cited by 23 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. OligoM-Cancer: A multidimensional information platform for deep phenotyping of heterogenous oligometastatic cancer;Computational and Structural Biotechnology Journal;2024-12

2. Performance of large language models in oral and maxillofacial surgery examinations;International Journal of Oral and Maxillofacial Surgery;2024-10

3. Analysis of Responses of GPT-4 V to the Japanese National Clinical Engineer Licensing Examination;Journal of Medical Systems;2024-09-11

4. Letter: Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations;Neurosurgery;2024-09-06

5. From GPT-3.5 to GPT-4.o: A Leap in AI’s Medical Exam Performance;Information;2024-09-05