Quality of Answers of Generative Large Language Models Versus Peer Users for Interpreting Laboratory Test Results for Lay Patients: Evaluation Study-Reference-Cited by-同舟云学术

Quality of Answers of Generative Large Language Models Versus Peer Users for Interpreting Laboratory Test Results for Lay Patients: Evaluation Study

Published:2024-04-17 Issue: Volume:26 Page:e56655
ISSN:1438-8871
Container-title:Journal of Medical Internet Research
language:en
Short-container-title:J Med Internet Res

Author:

He Zhe^ORCID,Bhasuran Balu^ORCID,Jin Qiao^ORCID,Tian Shubo^ORCID,Hanna Karim^ORCID,Shavor Cindy^ORCID,Arguello Lisbeth Garcia^ORCID,Murray Patrick^ORCID,Lu Zhiyong^ORCID

Abstract

Background Although patients have easy access to their electronic health records and laboratory test result data through patient portals, laboratory test results are often confusing and hard to understand. Many patients turn to web-based forums or question-and-answer (Q&A) sites to seek advice from their peers. The quality of answers from social Q&A sites on health-related questions varies significantly, and not all responses are accurate or reliable. Large language models (LLMs) such as ChatGPT have opened a promising avenue for patients to have their questions answered. Objective We aimed to assess the feasibility of using LLMs to generate relevant, accurate, helpful, and unharmful responses to laboratory test–related questions asked by patients and identify potential issues that can be mitigated using augmentation approaches. Methods We collected laboratory test result–related Q&A data from Yahoo! Answers and selected 53 Q&A pairs for this study. Using the LangChain framework and ChatGPT web portal, we generated responses to the 53 questions from 5 LLMs: GPT-4, GPT-3.5, LLaMA 2, MedAlpaca, and ORCA_mini. We assessed the similarity of their answers using standard Q&A similarity-based evaluation metrics, including Recall-Oriented Understudy for Gisting Evaluation, Bilingual Evaluation Understudy, Metric for Evaluation of Translation With Explicit Ordering, and Bidirectional Encoder Representations from Transformers Score. We used an LLM-based evaluator to judge whether a target model had higher quality in terms of relevance, correctness, helpfulness, and safety than the baseline model. We performed a manual evaluation with medical experts for all the responses to 7 selected questions on the same 4 aspects. Results Regarding the similarity of the responses from 4 LLMs; the GPT-4 output was used as the reference answer, the responses from GPT-3.5 were the most similar, followed by those from LLaMA 2, ORCA_mini, and MedAlpaca. Human answers from Yahoo data were scored the lowest and, thus, as the least similar to GPT-4–generated answers. The results of the win rate and medical expert evaluation both showed that GPT-4’s responses achieved better scores than all the other LLM responses and human responses on all 4 aspects (relevance, correctness, helpfulness, and safety). LLM responses occasionally also suffered from lack of interpretation in one’s medical context, incorrect statements, and lack of references. Conclusions By evaluating LLMs in generating responses to patients’ laboratory test result–related questions, we found that, compared to other 4 LLMs and human answers from a Q&A website, GPT-4’s responses were more accurate, helpful, relevant, and safer. There were cases in which GPT-4 responses were inaccurate and not individualized. We identified a number of ways to improve the quality of LLM responses, including prompt engineering, prompt augmentation, retrieval-augmented generation, and response evaluation.

Publisher

JMIR Publications Inc.

Reference52 articles.

1. Healthy people 2030: building a healthier future for allOffice of Disease Prevention and Health Promotion2023-05-09https://health.gov/healthypeople

2. NHE fact sheetCenters for Medicare & Medicaid Services2023-06-06https://tinyurl.com/yc4durw4

3. Prevention of chronic disease in the 21st century: elimination of the leading preventable causes of premature death and disability in the USA

4. Direct Release of Test Results to Patients Increases Patient Engagement and Utilization of Care

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. “Hospice Care Could Be a Compassionate Choice”: ChatGPT Responses to Questions About Decision Making in Advanced Cancer;Journal of Palliative Medicine;2024-09-12

2. Performance of Large Language Models in Patient Complaint Resolution: Web-Based Cross-Sectional Survey;Journal of Medical Internet Research;2024-08-09

3. Performance of ChatGPT Across Different Versions in Medical Licensing Examinations Worldwide: Systematic Review and Meta-Analysis;Journal of Medical Internet Research;2024-07-25

4. The professionalism of ChatGPT in the field of surgery: low or high level?;International Journal of Surgery;2024-05-09