Evaluating accuracy and reproducibility of ChatGPT responses to patient-based questions in Ophthalmology: An observational study-Reference-Cited by-同舟云学术

Evaluating accuracy and reproducibility of ChatGPT responses to patient-based questions in Ophthalmology: An observational study

Published:2024-08-09 Issue:32 Volume:103 Page:e39120
ISSN:0025-7974
Container-title:Medicine
language:en
Short-container-title:

Author:

Alqudah Asem A.¹^ORCID,Aleshawi Abdelwahab J.¹,Baker Mohammed¹,Alnajjar Zaina²,Ayasrah Ibrahim¹,Ta’ani Yaqoot¹,Al Salkhadi Mohammad¹,Aljawarneh Shaima’a¹

Affiliation:

1. Faculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan

2. Faculty of Medicine, Hashemite University, Zarqa, Jordan.

Abstract

Chat Generative Pre-Trained Transformer (ChatGPT) is an online large language model that appears to be a popular source of health information, as it can provide patients with answers in the form of human-like text, although the accuracy and safety of its responses are not evident. This study aims to evaluate the accuracy and reproducibility of ChatGPT responses to patients-based questions in ophthalmology. We collected 150 questions from the “Ask an ophthalmologist” page of the American Academy of Ophthalmology, which were reviewed and refined by two ophthalmologists for their eligibility. Each question was inputted into ChatGPT twice using the “new chat” option. The grading scale included the following: (1) comprehensive, (2) correct but inadequate, (3) some correct and some incorrect, and (4) completely incorrect. Totally, 117 questions were inputted into ChatGPT, which provided “comprehensive” responses to 70/117 (59.8%) of questions. Concerning reproducibility, it was defined as no difference in grading categories (1 and 2 vs 3 and 4) between the 2 responses for each question. ChatGPT provided reproducible responses to 91.5% of questions. This study shows moderate accuracy and reproducibility of ChatGPT responses to patients’ questions in ophthalmology. ChatGPT may be—after more modifications—a supplementary health information source, which should be used as an adjunct, but not a substitute, to medical advice. The reliability of ChatGPT should undergo more investigations.

Publisher

Ovid Technologies (Wolters Kluwer Health)

Reference19 articles.

1. Artificial intelligence in medicine.;Hamet;Metabolism,2017

2. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns.;Sallam;Healthcare,2023

3. ChatGPT and ophthalmology: exploring its potential with discharge summaries and operative notes.;Singh;Semin Ophthalmol,2023

4. Assessing the performance of ChatGPT in answering questions regarding cirrhosis and hepatocellular carcinoma.;Yeo;Clin Mol Hepatol,2023

5. Evaluate the accuracy of ChatGPT’s responses to diabetes questions and misconceptions.;Huang;J Transl Med,2023