PERFORMANCE OF CHAT GPT ON A TURKISH BOARD OF ORTHOPAEDİC SURGERY EXAMINATION (Preprint)-Reference-Cited by-同舟云学术

PERFORMANCE OF CHAT GPT ON A TURKISH BOARD OF ORTHOPAEDİC SURGERY EXAMINATION (Preprint)

Published:2024-07-15 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Ocak Bilgehan^ORCID

Abstract

UNSTRUCTURED

ABSTRACT Background: This study aimed to evaluate the success of the Chat GPT according to the Turkish Board of Orthopedic Surgery Examination Methods Among the written exam questions prepared by TOTEK between 2021 and 2023, questions asking visual information like that in the literature and canceled questions were not included, and all other questions were taken into consideration. The questions were divided into 19 categories according to topic. The questions were divided into 3 categories according to the methods of evaluating information: direct recall of information, ability to comment and ability to use information correctly. Questions were asked separately about the Chat GPT 3.5 and 4.0 artificial intelligence applications. All answers given were evaluated appropriately according to this grouping. Visual questions were not asked to the Chat GPT due to its inability to perceive visual questions. Only questions answered by the application with the correct choice and explanation were accepted as correct answers. Questions that were answered incorrectly by the Chat GPT were considered incorrect. Results We eliminated 300 visual questions in total and asked the remaining 265 multiple-choice questions about the Chat GPT. A total of 95 (35%) of 265 questions were answered correctly, and 169 (63%) were answered incorrectly. It was also seen that he could not answer 1 question. The exam success rate was greater for the Chat GPT group than for the control group, especially for the infection questions (67%). The descriptive findings are shown in Table 3, which shows that both artificial intelligence models can be effective at different levels on various issues, but predominantly, GPT 4 performs better. Conclusion Our study showed that although the Chat GPT could not reach the level of passing the Turkish Orthopedics and Traumatology Proficiency Exam, it could reach a certain level of accuracy. Software such as the Chat GPT needs to be developed and studied further to be useful for orthopedics and traumatology physicians, where the evaluation of radiological images and physical examination are very important.

Publisher

JMIR Publications Inc.

Reference14 articles.

1. Why People Use Chatbots

2. Artificial intelligence: Implications for the future of work

3. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models

4. Comparison of ChatGPT–3.5, ChatGPT-4, and Orthopaedic Resident Performance on Orthopaedic Assessment Examinations

5. Christian Terwiesch