Challenging the Chatbot: An Assessment of ChatGPT's Diagnoses and Recommendations for DBP Case Studies-Reference-Cited by-同舟云学术

Challenging the Chatbot: An Assessment of ChatGPT's Diagnoses and Recommendations for DBP Case Studies

Published:2024-02-09 Issue: Volume: Page:
ISSN:0196-206X
Container-title:Journal of Developmental & Behavioral Pediatrics
language:en
Short-container-title:J Dev Behav Pediatr

Author:

Kim Rachel,Margolis Alex,Barile Joe,Han Kyle,Kalash Saia,Papaioannou Helen,Krevskaya Anna,Milanaik Ruth

Abstract

Objective: Chat Generative Pretrained Transformer-3.5 (ChatGPT) is a publicly available and free artificial intelligence chatbot that logs billions of visits per day; parents may rely on such tools for developmental and behavioral medical consultations. The objective of this study was to determine how ChatGPT evaluates developmental and behavioral pediatrics (DBP) case studies and makes recommendations and diagnoses. Methods: ChatGPT was asked to list treatment recommendations and a diagnosis for each of 97 DBP case studies. A panel of 3 DBP physicians evaluated ChatGPT's diagnostic accuracy and scored treatment recommendations on accuracy (5-point Likert scale) and completeness (3-point Likert scale). Physicians also assessed whether ChatGPT's treatment plan correctly addressed cultural and ethical issues for relevant cases. Scores were analyzed using Python, and descriptive statistics were computed. Results: The DBP panel agreed with ChatGPT's diagnosis for 66.2% of the case reports. The mean accuracy score of ChatGPT's treatment plan was deemed by physicians to be 4.6 (between entirely correct and more correct than incorrect), and the mean completeness was 2.6 (between complete and adequate). Physicians agreed that ChatGPT addressed relevant cultural issues in 10 out of the 11 appropriate cases and the ethical issues in the single ethical case. Conclusion: While ChatGPT can generate a comprehensive and adequate list of recommendations, the diagnosis accuracy rate is still low. Physicians must advise caution to patients when using such online sources.

Publisher

Ovid Technologies (Wolters Kluwer Health)

Reference19 articles.

1. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models;Kung;PLOS Digit Health,2023

2. Diagnostic accuracy of differential-diagnosis lists generated by generative pretrained transformer 3 chatbot for clinical vignettes with common chief complaints: a pilot study;Hirosawa;Int J Environ Res Public Health,2023

3. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge;Kanjee;JAMA,2023

4. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum;Ayers;JAMA Intern Med,2023

5. An introduction to artificial intelligence in developmental and behavioral pediatrics;Aylward;J Dev Behav Pediatr,2023

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Reply;Journal of Developmental & Behavioral Pediatrics;2024-06-13

2. Perceptions of Machine Learning among Therapists Practicing Applied Behavior Analysis: A National Survey;Behavior Analysis in Practice;2024-05-23

3. ChatGPT's Diagnoses and Recommendations for Developmental-Behavioral Pediatrics Case Studies: Comment;Journal of Developmental & Behavioral Pediatrics;2024-05

4. Online Autism Diagnostic Evaluation: Its Rise, Promise, and Reasons for Caution;Journal of Developmental & Behavioral Pediatrics;2024-05

5. Future of ADHD Care: Evaluating the Efficacy of ChatGPT in Therapy Enhancement;Healthcare;2024-03-19