Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions-Reference-Cited by-同舟云学术

Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions

Published:2023-08-22 Issue:8 Volume:6 Page:e2330320
ISSN:2574-3805
Container-title:JAMA Network Open
language:en
Short-container-title:JAMA Netw Open

Author:

Bernstein Isaac A.¹,Zhang Youchen (Victor)¹,Govil Devendra¹,Majid Iyad¹,Chang Robert T.¹,Sun Yang¹,Shue Ann¹,Chou Jonathan C.²,Schehlein Emily³,Christopher Karen L.⁴,Groth Sylvia L.⁵,Ludwig Cassie¹,Wang Sophia Y.¹

Affiliation:

1. Department of Ophthalmology, Byers Eye Institute, Stanford University, Stanford, California

2. Department of Ophthalmology, Kaiser Permanente San Francisco, San Francisco, California

3. Brighton Vision Center, Brighton, Michigan

4. Department of Ophthalmology, University of Colorado School of Medicine, Aurora

5. Department of Ophthalmology and Visual Sciences, Vanderbilt Eye Institute, Nashville, Tennessee

Abstract

ImportanceLarge language models (LLMs) like ChatGPT appear capable of performing a variety of tasks, including answering patient eye care questions, but have not yet been evaluated in direct comparison with ophthalmologists. It remains unclear whether LLM-generated advice is accurate, appropriate, and safe for eye patients.ObjectiveTo evaluate the quality of ophthalmology advice generated by an LLM chatbot in comparison with ophthalmologist-written advice.Design, Setting, and ParticipantsThis cross-sectional study used deidentified data from an online medical forum, in which patient questions received responses written by American Academy of Ophthalmology (AAO)–affiliated ophthalmologists. A masked panel of 8 board-certified ophthalmologists were asked to distinguish between answers generated by the ChatGPT chatbot and human answers. Posts were dated between 2007 and 2016; data were accessed January 2023 and analysis was performed between March and May 2023.Main Outcomes and MeasuresIdentification of chatbot and human answers on a 4-point scale (likely or definitely artificial intelligence [AI] vs likely or definitely human) and evaluation of responses for presence of incorrect information, alignment with perceived consensus in the medical community, likelihood to cause harm, and extent of harm.ResultsA total of 200 pairs of user questions and answers by AAO-affiliated ophthalmologists were evaluated. The mean (SD) accuracy for distinguishing between AI and human responses was 61.3% (9.7%). Of 800 evaluations of chatbot-written answers, 168 answers (21.0%) were marked as human-written, while 517 of 800 human-written answers (64.6%) were marked as AI-written. Compared with human answers, chatbot answers were more frequently rated as probably or definitely written by AI (prevalence ratio [PR], 1.72; 95% CI, 1.52-1.93). The likelihood of chatbot answers containing incorrect or inappropriate material was comparable with human answers (PR, 0.92; 95% CI, 0.77-1.10), and did not differ from human answers in terms of likelihood of harm (PR, 0.84; 95% CI, 0.67-1.07) nor extent of harm (PR, 0.99; 95% CI, 0.80-1.22).Conclusions and RelevanceIn this cross-sectional study of human-written and AI-generated responses to 200 eye care questions from an online advice forum, a chatbot appeared capable of responding to long user-written eye health posts and largely generated appropriate responses that did not differ significantly from ophthalmologist-written responses in terms of incorrect information, likelihood of harm, extent of harm, or deviation from ophthalmologist community standards. Additional research is needed to assess patient attitudes toward LLM-augmented ophthalmologists vs fully autonomous AI content generation, to evaluate clarity and acceptability of LLM-generated answers from the patient perspective, to test the performance of LLMs in a greater variety of clinical contexts, and to determine an optimal manner of utilizing LLMs that is ethical and minimizes harm.

Publisher

American Medical Association (AMA)

Subject

General Medicine

Link

https://jamanetwork.com/journals/jamanetworkopen/articlepdf/2808557/bernstein_2023_oi_230872_1691786958.11886.pdf

Reference36 articles.

1. Length of stay prediction in neurosurgery with Russian GPT-3 language model compared to human expectations.;Danilov;Inform Technol Clin Care Public Health,2022

2. Medical image captioning via generative pretrained transformers.;Selivanov;Sci Rep,2023

3. Leveraging weak supervision to perform named entity recognition in electronic health records progress notes to identify the ophthalmology exam.;Wang;Int J Med Inform,2022

4. RadBERT: adapting transformer-based language models to radiology.;Yan;Radiol Artif Intell,2022

5. ChatGPT: the future of discharge summaries?;Patel;Lancet Digit Health,2023

Cited by 64 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Glaucoma diagnosis in the era of deep learning: A survey;Expert Systems with Applications;2024-12

2. Interpretation of Clinical Retinal Images Using an Artificial Intelligence Chatbot;Ophthalmology Science;2024-11

3. Human vs. AI: can ChatGPT improve tourism product descriptions?;Current Issues in Tourism;2024-09-13

4. Evaluating the effectiveness of large language models in patient education for conjunctivitis;British Journal of Ophthalmology;2024-08-30

5. Artificial intelligence applications in cataract and refractive surgeries;Current Opinion in Ophthalmology;2024-08-28