Unmasking and quantifying racial bias of large language models in medical report generation-Reference-Cited by-同舟云学术

Unmasking and quantifying racial bias of large language models in medical report generation

Published:2024-09-10 Issue:1 Volume:4 Page:
ISSN:2730-664X
Container-title:Communications Medicine
language:en
Short-container-title:Commun Med

Author:

Yang Yifan^ORCID,Liu Xiaoyu,Jin Qiao^ORCID,Huang Furong,Lu Zhiyong^ORCID

Abstract

Abstract Background Large language models like GPT-3.5-turbo and GPT-4 hold promise for healthcare professionals, but they may inadvertently inherit biases during their training, potentially affecting their utility in medical applications. Despite few attempts in the past, the precise impact and extent of these biases remain uncertain. Methods We use LLMs to generate responses that predict hospitalization, cost and mortality based on real patient cases. We manually examine the generated responses to identify biases. Results We find that these models tend to project higher costs and longer hospitalizations for white populations and exhibit optimistic views in challenging medical scenarios with much higher survival rates. These biases, which mirror real-world healthcare disparities, are evident in the generation of patient backgrounds, the association of specific diseases with certain racial and ethnic groups, and disparities in treatment recommendations, etc. Conclusions Our findings underscore the critical need for future research to address and mitigate biases in language models, especially in critical healthcare applications, to ensure fair and accurate outcomes for all patients.

Funder

U.S. Department of Health & Human Services | NIH | U.S. National Library of Medicine

Publisher

Springer Science and Business Media LLC

Link

https://www.nature.com/articles/s43856-024-00601-z.pdf

Reference20 articles.

1. Ouyang, L. et al. Training language models to follow instructions with human feedback. in Advances in Neural Information Processing Systems (eds. Koyejo, S. et al.) 35, 27730–27744 (Curran Associates, Inc., 2022).

2. OpenAI. GPT-4 Technical Report. Preprint at http://arxiv.org/abs/2303.08774 (2023).

3. Jin, Q., Wang, Z., Floudas, C. S., Sun, J. & Lu, Z. Matching Patients to Clinical Trials with Large Language Models. Preprint at https://doi.org/10.48550/arXiv.2307.15051 (2023).

4. Tian, S. et al. Opportunities and challenges for ChatGPT and large language models in biomedicine and health. Brief. Bioinform. 25, bbad493 (2024).

5. Zhuo, T. Y., Huang, Y., Chen, C. & Xing, Z. Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity. Preprint at https://doi.org/10.48550/arXiv.2301.12867 (2023).