Performance of ChatGPT on the Situational Judgement Test—A Professional Dilemmas–Based Examination for Doctors in the United Kingdom-Reference-Cited by-同舟云学术

Performance of ChatGPT on the Situational Judgement Test—A Professional Dilemmas–Based Examination for Doctors in the United Kingdom

Published:2023-08-07 Issue: Volume:9 Page:e48978
ISSN:2369-3762
Container-title:JMIR Medical Education
language:en
Short-container-title:JMIR Med Educ

Author:

Borchert Robin J^ORCID,Hickman Charlotte R^ORCID,Pepys Jack^ORCID,Sadler Timothy J^ORCID

Abstract

Background ChatGPT is a large language model that has performed well on professional examinations in the fields of medicine, law, and business. However, it is unclear how ChatGPT would perform on an examination assessing professionalism and situational judgement for doctors. Objective We evaluated the performance of ChatGPT on the Situational Judgement Test (SJT): a national examination taken by all final-year medical students in the United Kingdom. This examination is designed to assess attributes such as communication, teamwork, patient safety, prioritization skills, professionalism, and ethics. Methods All questions from the UK Foundation Programme Office’s (UKFPO’s) 2023 SJT practice examination were inputted into ChatGPT. For each question, ChatGPT’s answers and rationales were recorded and assessed on the basis of the official UK Foundation Programme Office scoring template. Questions were categorized into domains of Good Medical Practice on the basis of the domains referenced in the rationales provided in the scoring sheet. Questions without clear domain links were screened by reviewers and assigned one or multiple domains. ChatGPT's overall performance, as well as its performance across the domains of Good Medical Practice, was evaluated. Results Overall, ChatGPT performed well, scoring 76% on the SJT but scoring full marks on only a few questions (9%), which may reflect possible flaws in ChatGPT’s situational judgement or inconsistencies in the reasoning across questions (or both) in the examination itself. ChatGPT demonstrated consistent performance across the 4 outlined domains in Good Medical Practice for doctors. Conclusions Further research is needed to understand the potential applications of large language models, such as ChatGPT, in medical education for standardizing questions and providing consistent rationales for examinations assessing professionalism and ethics.

Publisher

JMIR Publications Inc.

Subject

Education

Reference17 articles.

1. ArXiv

2. ChatGPT Goes to Law School

3. TerwieschCWould Chat GPT Get a Wharton MBA? New White Paper By Christian TerwieschMack Institute for Innovation Management at the Wharton School, University of Pennsylvania20232023-07-26https://mackinstitute.wharton.upenn.edu/2023/would-chat-gpt3-get-a-wharton-mba-new-white-paper-by-christian-terwiesch/

4. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models

5. Performance of ChatGPT on UK Standardized Admission Tests: Insights From the BMAT, TMUA, LNAT, and TSA Examinations

Cited by 18 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Evaluating the interactions of Medical Doctors with chatbots based on large language models: Insights from a nationwide study in the Greek healthcare sector using ChatGPT;Computers in Human Behavior;2024-12

2. “Anything you can do, I can do”: Examining the use of ChatGPT in situational judgement tests for professional program admission;Journal of Vocational Behavior;2024-10

3. Integrating ChatGPT in Orthopedic Education for Medical Undergraduates: Randomized Controlled Trial;Journal of Medical Internet Research;2024-08-20

4. Generative artificial intelligence in healthcare: A scoping review on benefits, challenges and applications;International Journal of Medical Informatics;2024-08

5. Performance of ChatGPT Across Different Versions in Medical Licensing Examinations Worldwide: Systematic Review and Meta-Analysis;Journal of Medical Internet Research;2024-07-25