GPT-4 Multimodal Analysis on Ophthalmology Clinical Cases Including Text and Images-Reference-Cited by-同舟云学术

GPT-4 Multimodal Analysis on Ophthalmology Clinical Cases Including Text and Images

Published:2023-11-27 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Sorin Vera^ORCID,Kapelushnik Noa^ORCID,Hecht Idan^ORCID,Zloto Ofira^ORCID,Glicksberg Benjamin S.^ORCID,Bufman Hila^ORCID,Barash Yiftach^ORCID,Nadkarni Girish N.^ORCID,Klang Eyal^ORCID

Abstract

AbstractObjectiveRecent advancements in GPT-4 have enabled analysis of text with visual data. Diagnosis in ophthalmology is often based on ocular examinations and imaging, alongside the clinical context. The aim of this study was to evaluate the performance of multimodal GPT-4 (GPT-4V) in an integrated analysis of ocular images and clinical text.MethodsThis retrospective study included 40 patients seen in our institution with ocular pathologies. Cases were selected by a board certified ophthalmologist, to represent various pathologies and match the level for ophthalmology residents. We provided the model with each image, without and then with the clinical context. We also asked two non-ophthalmology physicians to write diagnoses for each image, without and then with the clinical context. Answers for both GPT-4V and the non-ophthalmologists were evaluated by two board-certified ophthalmologists. Performance accuracies were calculated and compared.ResultsGPT-4V provided the correct diagnosis in 19/40 (47.5%) cases based on images without clinical context, and in 27/40 (67.5%) cases when clinical context was provided. Non-ophthalmologists physicians provided the correct diagnoses in 24/40 (60.0%), and 23/40 (57.5%) of cases without clinical context, and in 29/40 (72.5%) and 27/40 (67.5%) with clinical context.ConclusionGPT-4V at its current stage is not yet suitable for clinical application in ophthalmology. Nonetheless, its ability to simultaneously analyze and integrate visual and textual data, and arrive at accurate clinical diagnoses in the majority of cases, is impressive. Multimodal large language models like GPT-4V have significant potential to advance both patient care and research in ophthalmology.

Publisher

Cold Spring Harbor Laboratory

Reference17 articles.

1. Nori H , King N , McKinney SM , Carignan D , Horvitz E (2023) Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:230313375

2. ChatGPT Surpasses 1000 Publications on PubMed: Envisioning the Road Ahead

3. The Use of ChatGPT to Assist in Diagnosing Glaucoma Based on Clinical Case Reports;Ophthalmology and Therapy,2023

4. What can GPT-4 do for Diagnosing Rare Eye Diseases? A Pilot Study;Ophthalmology and Therapy,2023

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. From vision to text: A comprehensive review of natural image captioning in medical diagnosis and radiology report generation;Medical Image Analysis;2024-10

2. Understanding natural language: Potential application of large language models to ophthalmology;Asia-Pacific Journal of Ophthalmology;2024-07

3. Deep Learning for Contrast Enhanced Mammography - a Systematic Review;2024-05-13

4. Advancing medical imaging with language models: featuring a spotlight on ChatGPT;Physics in Medicine & Biology;2024-05-03

5. Utility of artificial intelligence‐based large language models in ophthalmic care;Ophthalmic and Physiological Optics;2024-02-25