ICGA-GPT: report generation and question answering for indocyanine green angiography images-Reference-Cited by-同舟云学术

ICGA-GPT: report generation and question answering for indocyanine green angiography images

Published:2024-03-20 Issue: Volume: Page:bjo-2023-324446
ISSN:0007-1161
Container-title:British Journal of Ophthalmology
language:en
Short-container-title:Br J Ophthalmol

Author:

Chen Xiaolan,Zhang Weiyi,Zhao Ziwei,Xu Pusheng,Zheng Yingfeng^ORCID,Shi Danli^ORCID,He Mingguang

Abstract

BackgroundIndocyanine green angiography (ICGA) is vital for diagnosing chorioretinal diseases, but its interpretation and patient communication require extensive expertise and time-consuming efforts. We aim to develop a bilingual ICGA report generation and question-answering (QA) system.MethodsOur dataset comprised 213 129 ICGA images from 2919 participants. The system comprised two stages: image–text alignment for report generation by a multimodal transformer architecture, and large language model (LLM)-based QA with ICGA text reports and human-input questions. Performance was assessed using both qualitative metrics (including Bilingual Evaluation Understudy (BLEU), Consensus-based Image Description Evaluation (CIDEr), Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence (ROUGE-L), Semantic Propositional Image Caption Evaluation (SPICE), accuracy, sensitivity, specificity, precision and F1 score) and subjective evaluation by three experienced ophthalmologists using 5-point scales (5 refers to high quality).ResultsWe produced 8757 ICGA reports covering 39 disease-related conditions after bilingual translation (66.7% English, 33.3% Chinese). The ICGA-GPT model’s report generation performance was evaluated with BLEU scores (1–4) of 0.48, 0.44, 0.40 and 0.37; CIDEr of 0.82; ROUGE of 0.41 and SPICE of 0.18. For disease-based metrics, the average specificity, accuracy, precision, sensitivity and F1 score were 0.98, 0.94, 0.70, 0.68 and 0.64, respectively. Assessing the quality of 50 images (100 reports), three ophthalmologists achieved substantial agreement (kappa=0.723 for completeness, kappa=0.738 for accuracy), yielding scores from 3.20 to 3.55. In an interactive QA scenario involving 100 generated answers, the ophthalmologists provided scores of 4.24, 4.22 and 4.10, displaying good consistency (kappa=0.779).ConclusionThis pioneering study introduces the ICGA-GPT model for report generation and interactive QA for the first time, underscoring the potential of LLMs in assisting with automated ICGA image interpretation.

Funder

Start-up Fund for RAPs under the Strategic Hiring Scheme

Global STEM Professorship Scheme from HKSAR

Publisher

BMJ

Reference35 articles.

1. Translating color fundus photography to Indocyanine green angiography using deep-learning for age-related macular degeneration screening;Chen;NPJ Digit Med,2024

2. Indocyanine Green Angiography: A Perspective on Use in the Clinical Setting

3. Utility of a public-available artificial intelligence in diagnosis of Polypoidal Choroidal Vasculopathy;Yang;Graefes Arch Clin Exp Ophthalmol,2020

4. Polypoidal Choroidal Vasculopathy: an update on diagnosis and treatment;Sen;Clin Ophthalmol,2023

5. AI in health and medicine

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ChatFFA: An ophthalmic chat system for unified vision-language understanding and question answering for fundus fluorescein angiography;iScience;2024-07

2. Understanding natural language: Potential application of large language models to ophthalmology;Asia-Pacific Journal of Ophthalmology;2024-07

3. Unveiling the clinical incapabilities: a benchmarking study of GPT-4V(ision) for ophthalmic multimodal image analysis;British Journal of Ophthalmology;2024-05-24