Evaluating AI in medicine: a comparative analysis of expert and ChatGPT responses to colorectal cancer questions-Reference-Cited by-同舟云学术

Evaluating AI in medicine: a comparative analysis of expert and ChatGPT responses to colorectal cancer questions

Published:2024-02-03 Issue:1 Volume:14 Page:
ISSN:2045-2322
Container-title:Scientific Reports
language:en
Short-container-title:Sci Rep

Author:

Peng Wen,feng Yifei,Yao Cui,Zhang Sheng,Zhuo Han,Qiu Tianzhu,Zhang Yi,Tang Junwei,Gu Yanhong,Sun Yueming

Abstract

AbstractColorectal cancer (CRC) is a global health challenge, and patient education plays a crucial role in its early detection and treatment. Despite progress in AI technology, as exemplified by transformer-like models such as ChatGPT, there remains a lack of in-depth understanding of their efficacy for medical purposes. We aimed to assess the proficiency of ChatGPT in the field of popular science, specifically in answering questions related to CRC diagnosis and treatment, using the book “Colorectal Cancer: Your Questions Answered” as a reference. In general, 131 valid questions from the book were manually input into ChatGPT. Responses were evaluated by clinical physicians in the relevant fields based on comprehensiveness and accuracy of information, and scores were standardized for comparison. Not surprisingly, ChatGPT showed high reproducibility in its responses, with high uniformity in comprehensiveness, accuracy, and final scores. However, the mean scores of ChatGPT’s responses were significantly lower than the benchmarks, indicating it has not reached an expert level of competence in CRC. While it could provide accurate information, it lacked in comprehensiveness. Notably, ChatGPT performed well in domains of radiation therapy, interventional therapy, stoma care, venous care, and pain control, almost rivaling the benchmarks, but fell short in basic information, surgery, and internal medicine domains. While ChatGPT demonstrated promise in specific domains, its general efficiency in providing CRC information falls short of expert standards, indicating the need for further advancements and improvements in AI technology for patient education in healthcare.

Publisher

Springer Science and Business Media LLC

Link

https://www.nature.com/articles/s41598-024-52853-3.pdf

Reference20 articles.

1. Bando, H., Ohtsu, A. & Yoshino, T. Therapeutic landscape and future direction of metastatic colorectal cancer. Nat. Rev. Gastroenterol. Hepatol. 20(5), 306–322 (2023).

2. Li, Q. et al. Colorectal cancer burden, trends and risk factors in China: A review and comparison with the United States. Chin. J. Cancer Res. 34(5), 483–495 (2022).

3. Kruk, M. E. et al. High-quality health systems in the Sustainable Development Goals era: Time for a revolution. Lancet Glob. Health 6(11), e1196–e1252 (2018).

4. Loomans-Kropp, H. A. & Umar, A. Cancer prevention and screening: The next step in the era of precision medicine. NPJ Precis. Oncol. 3, 3 (2019).

5. Walter, F., Webster, A., Scott, S. & Emery, J. The Andersen Model of Total Patient Delay: A systematic review of its application in cancer diagnosis. J. Health Serv. Res. Policy 17(2), 110–118 (2012).

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A new era in medical information: ChatGPT outperforms medical information provided by online information sheets about congenital malformations;Journal of Pediatric Surgery;2024-09

2. Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review;2024-06-27

3. Toward Clinical Generative AI: Conceptual Framework;JMIR AI;2024-06-07

4. Evaluating the Accuracy, Comprehensiveness, and Validity of ChatGPT Compared to Evidence-Based Sources Regarding Common Surgical Conditions: Surgeons’ Perspectives;The American Surgeon™;2024-05-25