Comparison of the diagnostic accuracy among GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists in musculoskeletal radiology-Reference-Cited by-同舟云学术

Comparison of the diagnostic accuracy among GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists in musculoskeletal radiology

Published:2023-12-09 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Horiuchi Daisuke^ORCID,Tatekawa Hiroyuki^ORCID,Oura Tatsushi^ORCID,Shimono Taro^ORCID,Walston Shannon L^ORCID,Takita Hirotaka^ORCID,Matsushita Shu^ORCID,Mitsuyama Yasuhito^ORCID,Miki Yukio^ORCID,Ueda Daiju^ORCID

Abstract

AbstractObjectiveTo compare the diagnostic accuracy of Generative Pre-trained Transformer (GPT)-4 based ChatGPT, GPT-4 with vision (GPT-4V) based ChatGPT, and radiologists in musculoskeletal radiology.Materials and MethodsWe included 106 “Test Yourself” cases fromSkeletal Radiologybetween January 2014 and September 2023. We input the medical history and imaging findings into GPT-4 based ChatGPT and the medical history and images into GPT-4V based ChatGPT, then both generated a diagnosis for each case. Two radiologists (a radiology resident and a board-certified radiologist) independently provided diagnoses for all cases. The diagnostic accuracy rates were determined based on the published ground truth. Chi-square tests were performed to compare the diagnostic accuracy of GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists.ResultsGPT-4 based ChatGPT significantly outperformed GPT-4V based ChatGPT (p< 0.001) with accuracy rates of 43% (46/106) and 8% (9/106), respectively. The radiology resident and the board-certified radiologist achieved accuracy rates of 41% (43/106) and 53% (56/106). The diagnostic accuracy of GPT-4 based ChatGPT was comparable to that of the radiology resident but was lower than that of the board-certified radiologist, although the differences were not significant (p= 0.78 and 0.22, respectively). The diagnostic accuracy of GPT-4V based ChatGPT was significantly lower than those of both radiologists (p< 0.001 and < 0.001, respectively).ConclusionGPT-4 based ChatGPT demonstrated significantly higher diagnostic accuracy than GPT-4V based ChatGPT. While GPT-4 based ChatGPT’s diagnostic performance was comparable to radiology residents, it did not reach the performance level of board-certified radiologists in musculoskeletal radiology.

Publisher

Cold Spring Harbor Laboratory

Reference32 articles.

1. OpenAI. GPT-4 technical report. arXiv [cs.CL]. 2023. http://arxiv.org/abs/2303.08774

2. Brown TB , Mann B , Ryder N , et al. Language models are few-shot learners. arXiv [cs.CL]. 2020. https://arxiv.org/abs/2005.14165

3. Bubeck S , Chandrasekaran V , Eldan R , et al. Sparks of artificial general intelligence: early experiments with GPT-4. arXiv [cs.CL]. 2023. http://arxiv.org/abs/2303.12712

4. Eloundou T , Manning S , Mishkin P , et al. GPTs are GPTs: an early look at the labor market impact potential of large language models. arXiv [econ.GN]. 2023. http://arxiv.org/abs/2303.10130

5. OpenAI. GPT-4V(ision) system card. [Internet] 2023 Sep 25 [cited 2023 October 13]; Available from: https://openai.com/research/gpt-4v-system-card.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Diagnostic Performance of Generative AI and Physicians: A Systematic Review and Meta-Analysis;2024-01-22