Evaluation of ChatGPT in providing appropriate fracture prevention recommendations and medical science question responses: A quantitative research-Reference-Cited by-同舟云学术

Evaluation of ChatGPT in providing appropriate fracture prevention recommendations and medical science question responses: A quantitative research

Published:2024-03-15 Issue:11 Volume:103 Page:e37458
ISSN:0025-7974
Container-title:Medicine
language:en
Short-container-title:

Author:

Meng Jiahao¹,Zhang Ziyi²,Tang Hang¹,Xiao Yifan¹,Liu Pan¹,Gao Shuguang¹³,He Miao²^ORCID

Affiliation:

1. Department of Orthopaedics, Xiangya Hospital, Central South University, #87 Xiangya Road, Changsha, Hunan, China

2. Department of Neurology, The Second Xiangya Hospital, Central South University, Changsha, Hunan, China

3. National Clinical Research Center of Geriatric Disorders, Xiangya Hospital, Central South University, Changsha, Hunan, China.

Abstract

Currently, there are limited studies assessing ChatGPT ability to provide appropriate responses to medical questions. Our study aims to evaluate ChatGPT adequacy in responding to questions regarding osteoporotic fracture prevention and medical science. We created a list of 25 questions based on the guidelines and our clinical experience. Additionally, we included 11 medical science questions from the journal Science. Three patients, 3 non-medical professionals, 3 specialist doctor and 3 scientists were involved to evaluate the accuracy and appropriateness of responses by ChatGPT3.5 on October 2, 2023. To simulate a consultation, an inquirer (either a patient or non-medical professional) would send their questions to a consultant (specialist doctor or scientist) via a website. The consultant would forward the questions to ChatGPT for answers, which would then be evaluated for accuracy and appropriateness by the consultant before being sent back to the inquirer via the website for further review. The primary outcome is the appropriate, inappropriate, and unreliable rate of ChatGPT responses as evaluated separately by the inquirer and consultant groups. Compared to orthopedic clinicians, the patients’ rating on the appropriateness of ChatGPT responses to the questions about osteoporotic fracture prevention was slightly higher, although the difference was not statistically significant (88% vs 80%, P = .70). For medical science questions, non-medical professionals and medical scientists rated similarly. In addition, the experts’ ratings on the appropriateness of ChatGPT responses to osteoporotic fracture prevention and to medical science questions were comparable. On the other hand, the patients perceived that the appropriateness of ChatGPT responses to osteoporotic fracture prevention questions was slightly higher than that to medical science questions (88% vs 72·7%, P = .34). ChatGPT is capable of providing comparable and appropriate responses to medical science questions, as well as to fracture prevention related issues. Both the inquirers seeking advice and the consultants providing advice recognize ChatGPT expertise in these areas.

Publisher

Ovid Technologies (Wolters Kluwer Health)

Reference8 articles.

1. ChatGPT: friend or foe?;Lancet Digit Health,2023

2. ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports.;Jeblick;Eur Radiol,2023

3. ChatGPT: the future of discharge summaries?;Patel;Lancet Digit Health,2023

4. Appropriateness of cardiovascular disease prevention recommendations obtained from a popular online chat-based artificial intelligence model.;Sarraju;JAMA,2023

5. UK clinical guideline for the prevention and treatment of osteoporosis.;Gregson;Arch Osteoporos,2022

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The Role of Artificial Intelligence in the Primary Prevention of Common Musculoskeletal Diseases;Cureus;2024-07-25