Performance of generative pre-trained transformers (GPTs) in Certification Examination of the College of Family Physicians of Canada-Reference-Cited by-同舟云学术

Performance of generative pre-trained transformers (GPTs) in Certification Examination of the College of Family Physicians of Canada

Published:2024-05 Issue:Suppl 1 Volume:12 Page:e002626
ISSN:2305-6983
Container-title:Family Medicine and Community Health
language:en
Short-container-title:Fam Med Com Health

Author:

Mousavi Mehdi^ORCID,Shafiee Shabnam,Harley Jason M,Cheung Jackie Chi Kit,Abbasgholizadeh Rahimi Samira^ORCID

Abstract

IntroductionThe application of large language models such as generative pre-trained transformers (GPTs) has been promising in medical education, and its performance has been tested for different medical exams. This study aims to assess the performance of GPTs in responding to a set of sample questions of short-answer management problems (SAMPs) from the certification exam of the College of Family Physicians of Canada (CFPC).MethodBetween August 8th and 25th, 2023, we used GPT-3.5 and GPT-4 in five rounds to answer a sample of 77 SAMPs questions from the CFPC website. Two independent certified family physician reviewers scored AI-generated responses twice: first, according to the CFPC answer key (ie, CFPC score), and second, based on their knowledge and other references (ie, Reviews’ score). An ordinal logistic generalised estimating equations (GEE) model was applied to analyse repeated measures across the five rounds.ResultAccording to the CFPC answer key, 607 (73.6%) lines of answers by GPT-3.5 and 691 (81%) by GPT-4 were deemed accurate. Reviewer’s scoring suggested that about 84% of the lines of answers provided by GPT-3.5 and 93% of GPT-4 were correct. The GEE analysis confirmed that over five rounds, the likelihood of achieving a higher CFPC Score Percentage for GPT-4 was 2.31 times more than GPT-3.5 (OR: 2.31; 95% CI: 1.53 to 3.47; p<0.001). Similarly, the Reviewers’ Score percentage for responses provided by GPT-4 over 5 rounds were 2.23 times more likely to exceed those of GPT-3.5 (OR: 2.23; 95% CI: 1.22 to 4.06; p=0.009). Running the GPTs after a one week interval, regeneration of the prompt or using or not using the prompt did not significantly change the CFPC score percentage.ConclusionIn our study, we used GPT-3.5 and GPT-4 to answer complex, open-ended sample questions of the CFPC exam and showed that more than 70% of the answers were accurate, and GPT-4 outperformed GPT-3.5 in responding to the questions. Large language models such as GPTs seem promising for assisting candidates of the CFPC exam by providing potential answers. However, their use for family medicine education and exam preparation needs further studies.

Publisher

BMJ

Reference24 articles.

1. OpenAI . Models: OpenAI. 2023. Available: https://beta.openai.com/docs/models

2. Benoit JRA . ChatGPT for clinical vignette generation, revision, and evaluation. medRxiv 2023;2023. doi:10.1101/2023.02.04.23285478

3. ChatGPT - reshaping medical education and clinical management;Khan;Pak J Med Sci,2023

4. Diagnostic accuracy of differential-diagnosis lists generated by generative pretrained transformer 3 Chatbot for clinical vignettes with common chief complaints: a pilot study;Hirosawa;Int J Environ Res Public Health,2023

5. The next paradigm shift? ChatGPT, artificial intelligence, and medical education;Wang;Medical Teacher,2023

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Analysis of Responses of GPT-4 V to the Japanese National Clinical Engineer Licensing Examination;Journal of Medical Systems;2024-09-11

2. Potential of ChatGPT to Pass the Japanese Medical and Healthcare Professional National Licenses: A Literature Review;Cureus;2024-08-06