Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study-Reference-Cited by-同舟云学术

Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study

Published:2024-01-22 Issue: Volume:26 Page:e51926
ISSN:1438-8871
Container-title:Journal of Medical Internet Research
language:en
Short-container-title:J Med Internet Res

Author:

Liu Xiaocong^ORCID,Wu Jiageng^ORCID,Shao An^ORCID,Shen Wenyue^ORCID,Ye Panpan^ORCID,Wang Yao^ORCID,Ye Juan^ORCID,Jin Kai^ORCID,Yang Jie^ORCID

Abstract

Background Benefiting from rich knowledge and the exceptional ability to understand text, large language models like ChatGPT have shown great potential in English clinical environments. However, the performance of ChatGPT in non-English clinical settings, as well as its reasoning, have not been explored in depth. Objective This study aimed to evaluate ChatGPT’s diagnostic performance and inference abilities for retinal vascular diseases in a non-English clinical environment. Methods In this cross-sectional study, we collected 1226 fundus fluorescein angiography reports and corresponding diagnoses written in Chinese and tested ChatGPT with 4 prompting strategies (direct diagnosis or diagnosis with a step-by-step reasoning process and in Chinese or English). Results Compared with ChatGPT using Chinese prompts for direct diagnosis that achieved an F1-score of 70.47%, ChatGPT using English prompts for direct diagnosis achieved the best diagnostic performance (80.05%), which was inferior to ophthalmologists (89.35%) but close to ophthalmologist interns (82.69%). As for its inference abilities, although ChatGPT can derive a reasoning process with a low error rate (0.4 per report) for both Chinese and English prompts, ophthalmologists identified that the latter brought more reasoning steps with less incompleteness (44.31%), misinformation (1.96%), and hallucinations (0.59%) (all P<.001). Also, analysis of the robustness of ChatGPT with different language prompts indicated significant differences in the recall (P=.03) and F1-score (P=.04) between Chinese and English prompts. In short, when prompted in English, ChatGPT exhibited enhanced diagnostic and inference capabilities for retinal vascular disease classification based on Chinese fundus fluorescein angiography reports. Conclusions ChatGPT can serve as a helpful medical assistant to provide diagnosis in non-English clinical environments, but there are still performance gaps, language disparities, and errors compared to professionals, which demonstrate the potential limitations and the need to continually explore more robust large language models in ophthalmology practice.

Publisher

JMIR Publications Inc.

Subject

Health Informatics

Reference46 articles.

1. Trends in prevalence of blindness and distance and near vision impairment over 30 years: an analysis for the Global Burden of Disease Study

2. Nanoengineering of therapeutics for retinal vascular disease

4. Automatic interpretation and clinical evaluation for fundus fluorescein angiography images of diabetic retinopathy patients by deep learning

5. Multi-label classification of retinal lesions in diabetic retinopathy for automatic analysis of fundus fluorescein angiography based on deep learning

Cited by 11 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The performance of OpenAI ChatGPT-4 and Google Gemini in virology multiple-choice questions: a comparative analysis of English and Arabic responses;BMC Research Notes;2024-09-03

2. Ethical considerations for large language models in ophthalmology;Current Opinion in Ophthalmology;2024-08-27

3. Language discrepancies in the performance of generative artificial intelligence models: an examination of infectious disease queries in English and Arabic;BMC Infectious Diseases;2024-08-08

4. Understanding natural language: Potential application of large language models to ophthalmology;Asia-Pacific Journal of Ophthalmology;2024-07

5. Vision of the future: large language models in ophthalmology;Current Opinion in Ophthalmology;2024-05-30