Race, Sex, and Age Disparities in the Performance of ECG Deep Learning Models Predicting Heart Failure-Reference-Cited by-同舟云学术

Race, Sex, and Age Disparities in the Performance of ECG Deep Learning Models Predicting Heart Failure

Published:2024-01 Issue:1 Volume:17 Page:
ISSN:1941-3289
Container-title:Circulation: Heart Failure
language:en
Short-container-title:Circ: Heart Failure

Author:

Kaur Dhamanpreet¹^ORCID,Hughes J. Weston¹,Rogers Albert J.¹^ORCID,Kang Guson²,Narayan Sanjiv M.¹^ORCID,Ashley Euan A.¹^ORCID,Perez Marco V.¹^ORCID

Affiliation:

1. Cardiovascular Institute, Stanford University, CA (D.K., J.W.H., A.J.R., S.M.N., E.A.A., M.V.P.).

2. VA Palo Alto Health Care System, Cardiovascular Medicine, CA (G.K.).

Abstract

BACKGROUND: Deep learning models may combat widening racial disparities in heart failure outcomes through early identification of individuals at high risk. However, demographic biases in the performance of these models have not been well-studied. METHODS: This retrospective analysis used 12-lead ECGs taken between 2008 and 2018 from 326 518 patient encounters referred for standard clinical indications to Stanford Hospital. The primary model was a convolutional neural network model trained to predict incident heart failure within 5 years. Biases were evaluated on the testing set (160 312 ECGs) using the area under the receiver operating characteristic curve, stratified across the protected attributes of race, ethnicity, age, and sex. RESULTS: There were 59 817 cases of incident heart failure observed within 5 years of ECG collection. The performance of the primary model declined with age. There were no significant differences observed between racial groups overall. However, the primary model performed significantly worse in Black patients aged 0 to 40 years compared with all other racial groups in this age group, with differences most pronounced among young Black women. Disparities in model performance did not improve with the integration of race, ethnicity, sex, and age into model architecture, by training separate models for each racial group, or by providing the model with a data set of equal racial representation. Using probability thresholds individualized for race, age, and sex offered substantial improvements in F1 scores. CONCLUSIONS: The biases found in this study warrant caution against perpetuating disparities through the development of machine learning tools for the prognosis and management of heart failure. Customizing the application of these models by using probability thresholds individualized by race, ethnicity, age, and sex may offer an avenue to mitigate existing algorithmic disparities.

Publisher

Ovid Technologies (Wolters Kluwer Health)

Subject

Cardiology and Cardiovascular Medicine

Reference38 articles.

1. Heart Disease and Stroke Statistics—2020 Update: A Report From the American Heart Association

2. Heart failure in primary care: prevalence related to age and comorbidity

3. Epidemiology of heart failure

4. Disparities in Cardiovascular Mortality Related to Heart Failure in the United States

5. Disparity in the Setting of Incident Heart Failure Diagnosis

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Artificial Intelligence to Promote Racial and Ethnic Cardiovascular Health Equity;Current Cardiovascular Risk Reports;2024-08-20

2. Mitigating the risk of artificial intelligence bias in cardiovascular care;The Lancet Digital Health;2024-08

3. Understanding AI bias in clinical practice;Heart Rhythm;2024-08

4. Diagnostic and Prognostic Electrocardiogram-Based Models for Rapid Clinical Applications;Canadian Journal of Cardiology;2024-07

5. Simple models vs. deep learning in detecting low ejection fraction from the electrocardiogram;European Heart Journal - Digital Health;2024-04-25