Performance of ChatGPT on Chinese Master’s Degree Entrance Examination in Clinical Medicine-Reference-Cited by-同舟云学术

Performance of ChatGPT on Chinese Master’s Degree Entrance Examination in Clinical Medicine

Published:2024-04-04 Issue:4 Volume:19 Page:e0301702
ISSN:1932-6203
Container-title:PLOS ONE
language:en
Short-container-title:PLoS ONE

Author:

Li Ke-Cheng,Bu Zhi-Jun,Shahjalal Md.^ORCID,He Bai-Xiang,Zhuang Zi-Fan,Li Chen,Liu Jian-Ping,Wang Bin,Liu Zhao-Lan^ORCID

Abstract

Background ChatGPT is a large language model designed to generate responses based on a contextual understanding of user queries and requests. This study utilised the entrance examination for the Master of Clinical Medicine in Traditional Chinese Medicine to assesses the reliability and practicality of ChatGPT within the domain of medical education. Methods We selected 330 single and multiple-choice questions from the 2021 and 2022 Chinese Master of Clinical Medicine comprehensive examinations, which did not include any images or tables. To ensure the test’s accuracy and authenticity, we preserved the original format of the query and alternative test texts, without any modifications or explanations. Results Both ChatGPT3.5 and GPT-4 attained average scores surpassing the admission threshold. Noteworthy is that ChatGPT achieved the highest score in the Medical Humanities section, boasting a correct rate of 93.75%. However, it is worth noting that ChatGPT3.5 exhibited the lowest accuracy percentage of 37.5% in the Pathology division, while GPT-4 also displayed a relatively lower correctness percentage of 60.23% in the Biochemistry section. An analysis of sub-questions revealed that ChatGPT demonstrates superior performance in handling single-choice questions but performs poorly in multiple-choice questions. Conclusion ChatGPT exhibits a degree of medical knowledge and the capacity to aid in diagnosing and treating diseases. Nevertheless, enhancements are warranted to address its accuracy and reliability limitations. Imperatively, rigorous evaluation and oversight must accompany its utilization, accompanied by proactive measures to surmount prevailing constraints.

Funder

National Natural Science Foundation of China

Reserve Discipline Leader Funding of Beijing University of Chinese Medicine

Publisher

Public Library of Science (PLoS)

Reference19 articles.

1. OpenAI R. Gpt-4 technical report. arxiv 2303.08774. View in Article, 2023, 2.

2. Role of Chat GPT in Public Health;SS Biswas;Ann Biomed Eng,2023

3. ChatGPT: Jack of all trades, master of none[J];J Kocoń;Information Fusion,2023

4. Koubaa, A. GPT-4 vs. GPT-3.5: A Concise Showdown. TechRxiv.2023.

5. An era of ChatGPT as a significant futuristic support tool: A study on features, abilities, and challenges[J];A Haleem;BenchCouncil transactions on benchmarks, standards and evaluations,2022

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Assessing ChatGPT-4's Proficiency in English College Entrance Examinations Using Web Raschonline: A Comparative Study (Preprint);2024-07-19