Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced-Reference-Cited by-同舟云学术

Reliability of large language models for advanced head and neck malignancies management: a comparison between ChatGPT 4 and Gemini Advanced

Published:2024-05-25 Issue:9 Volume:281 Page:5001-5006
ISSN:0937-4477
Container-title:European Archives of Oto-Rhino-Laryngology
language:en
Short-container-title:Eur Arch Otorhinolaryngol

Author:

Lorenzi Andrea,Pugliese Giorgia^ORCID,Maniaci Antonino,Lechien Jerome R.,Allevi Fabiana,Boscolo-Rizzo Paolo,Vaira Luigi Angelo,Saibene Alberto Maria

Abstract

Abstract Purpose This study evaluates the efficacy of two advanced Large Language Models (LLMs), OpenAI’s ChatGPT 4 and Google’s Gemini Advanced, in providing treatment recommendations for head and neck oncology cases. The aim is to assess their utility in supporting multidisciplinary oncological evaluations and decision-making processes. Methods This comparative analysis examined the responses of ChatGPT 4 and Gemini Advanced to five hypothetical cases of head and neck cancer, each representing a different anatomical subsite. The responses were evaluated against the latest National Comprehensive Cancer Network (NCCN) guidelines by two blinded panels using the total disagreement score (TDS) and the artificial intelligence performance instrument (AIPI). Statistical assessments were performed using the Wilcoxon signed-rank test and the Friedman test. Results Both LLMs produced relevant treatment recommendations with ChatGPT 4 generally outperforming Gemini Advanced regarding adherence to guidelines and comprehensive treatment planning. ChatGPT 4 showed higher AIPI scores (median 3 [2–4]) compared to Gemini Advanced (median 2 [2–3]), indicating better overall performance. Notably, inconsistencies were observed in the management of induction chemotherapy and surgical decisions, such as neck dissection. Conclusions While both LLMs demonstrated the potential to aid in the multidisciplinary management of head and neck oncology, discrepancies in certain critical areas highlight the need for further refinement. The study supports the growing role of AI in enhancing clinical decision-making but also emphasizes the necessity for continuous updates and validation against current clinical standards to integrate AI into healthcare practices fully.

Funder

Università degli Studi di Milano

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s00405-024-08746-2.pdf

Reference12 articles.

1. Liu S et al (2023) Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J Am Med Inform Assoc 30:1237–1245

2. Marchi F, Bellini E, Iandelli A, Sampieri C, Peretti G (2024) Exploring the landscape of AI-assisted decision-making in head and neck cancer treatment: a comparative analysis of NCCN guidelines and ChatGPT responses. Eur Arch Otorhinolaryngol 281:2123–2136

3. Sarma G, Kashyap H, Medhi PP (2024) ChatGPT in head and neck oncology-opportunities and challenges. Indian J Otolaryngol Head Neck Surg 76:1425–1429

4. Saibene AM et al (2024) Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation. Eur Arch Otorhinolaryngol 281:1835–1841

5. Vaira LA et al (2023) Accuracy of ChatGPT-generated information on head and neck and oromaxillofacial surgery: a multicenter collaborative analysis. Otolaryngol Head Neck Surg. https://doi.org/10.1002/ohn.489

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Evaluation of Vertigo-Related Information from Artificial Intelligence Chatbot;2024-09-02