Artificial intelligence derived large language model in decision-making process in uveitis-Reference-Cited by-同舟云学术

Artificial intelligence derived large language model in decision-making process in uveitis

Published:2024-09-11 Issue:1 Volume:10 Page:
ISSN:2056-9920
Container-title:International Journal of Retina and Vitreous
language:en
Short-container-title:Int J Retin Vitr

Author:

Schumacher Inès,Bühler Virginie Manuela Marie,Jaggi Damian,Roth Janice

Abstract

Abstract Background Uveitis is the ophthalmic subfield dealing with a broad range of intraocular inflammatory diseases. With the raising importance of LLM such as ChatGPT and their potential use in the medical field, this research explores the strengths and weaknesses of its applicability in the subfield of uveitis. Methods A series of highly clinically relevant questions were asked three consecutive times (attempts 1, 2 and 3) of the LLM regarding current uveitis cases. The answers were classified on whether they were accurate and sufficient, partially accurate and sufficient or inaccurate and insufficient. Statistical analysis included descriptive analysis, normality distribution, non-parametric test and reliability tests. References were checked for their correctness in different medical databases. Results The data showed non-normal distribution. Data between subgroups (attempts 1, 2 and 3) was comparable (Kruskal-Wallis H test, p-value = 0.7338). There was a moderate agreement between attempt 1 and attempt 2 (Cohen’s kappa, ĸ = 0.5172) as well as between attempt 2 and attempt 3 (Cohen’s kappa, ĸ = 0.4913). There was a fair agreement between attempt 1 and attempt 3 (Cohen’s kappa, ĸ = 0.3647). The average agreement was moderate (Cohen’s kappa, ĸ = 0.4577). Between the three attempts together, there was a moderate agreement (Fleiss’ kappa, ĸ = 0.4534). A total of 52 references were generated by the LLM. 22 references (42.3%) were found to be accurate and correctly cited. Another 22 references (42.3%) could not be located in any of the searched databases. The remaining 8 references (15.4%) were found to exist, but were either misinterpreted or incorrectly cited by the LLM. Conclusion Our results demonstrate the significant potential of LLMs in uveitis. However, their implementation requires rigorous training and comprehensive testing for specific medical tasks. We also found out that the references made by ChatGPT 4.o were in most cases incorrect. LLMs are likely to become invaluable tools in shaping the future of ophthalmology, enhancing clinical decision-making and patient care.

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1186/s40942-024-00581-1.pdf

Reference13 articles.

1. Anguita R, Downie C, Ferro Desideri L, Sagoo MS. Assessing large language models' accuracy in providing patient support for choroidal melanoma. Eye (Lond). 2024. https://doi.org/10.1038/s41433-024-03231-w.

2. Anguita R, Makuloluwa A, Hind J, Wickham L. Large language models in vitreoretinal surgery. Eye. 2024;38(4):809–10. https://doi.org/10.1038/s41433-023-02751-1.

3. Ferro Desideri L, Roth J, Zinkernagel M, Anguita R. Application and accuracy of artificial intelligence-derived large language models in patients with age related macular degeneration. Int J Retina Vitr. 2023;9(1):71. https://doi.org/10.1186/s40942-023-00511-7.

4. Rojas-Carabali W, Sen A, Agarwal A, et al. Chatbots Vs. Human experts: evaluating diagnostic performance of Chatbots in Uveitis and the perspectives on AI adoption in Ophthalmology. Ocul Immunol Inflamm. 2023;0(0):1–8. https://doi.org/10.1080/09273948.2023.2266730.

5. Jacquot R, Sève P, Jackson TL, Wang T, Duclos A, Stanescu-Segall D. Diagnosis, classification, and Assessment of the underlying etiology of Uveitis by Artificial Intelligence: a systematic review. J Clin Med. 2023;12(11):3746. https://doi.org/10.3390/jcm12113746.