Machine Learning–Based Prediction Models for Different Clinical Risks in Different Hospitals: Evaluation of Live Performance-Reference-Cited by-同舟云学术

Machine Learning–Based Prediction Models for Different Clinical Risks in Different Hospitals: Evaluation of Live Performance

Published:2022-06-07 Issue:6 Volume:24 Page:e34295
ISSN:1438-8871
Container-title:Journal of Medical Internet Research
language:en
Short-container-title:J Med Internet Res

Author:

Sun Hong^ORCID,Depraetere Kristof^ORCID,Meesseman Laurent^ORCID,Cabanillas Silva Patricia^ORCID,Szymanowsky Ralph^ORCID,Fliegenschmidt Janis^ORCID,Hulde Nikolai^ORCID,von Dossow Vera^ORCID,Vanbiervliet Martijn^ORCID,De Baerdemaeker Jos^ORCID,Roccaro-Waldmeyer Diana M^ORCID,Stieg Jörg^ORCID,Domínguez Hidalgo Manuel^ORCID,Dahlweid Fried-Michael^ORCID

Abstract

Background Machine learning algorithms are currently used in a wide array of clinical domains to produce models that can predict clinical risk events. Most models are developed and evaluated with retrospective data, very few are evaluated in a clinical workflow, and even fewer report performances in different hospitals. In this study, we provide detailed evaluations of clinical risk prediction models in live clinical workflows for three different use cases in three different hospitals. Objective The main objective of this study was to evaluate clinical risk prediction models in live clinical workflows and compare their performance in these setting with their performance when using retrospective data. We also aimed at generalizing the results by applying our investigation to three different use cases in three different hospitals. Methods We trained clinical risk prediction models for three use cases (ie, delirium, sepsis, and acute kidney injury) in three different hospitals with retrospective data. We used machine learning and, specifically, deep learning to train models that were based on the Transformer model. The models were trained using a calibration tool that is common for all hospitals and use cases. The models had a common design but were calibrated using each hospital’s specific data. The models were deployed in these three hospitals and used in daily clinical practice. The predictions made by these models were logged and correlated with the diagnosis at discharge. We compared their performance with evaluations on retrospective data and conducted cross-hospital evaluations. Results The performance of the prediction models with data from live clinical workflows was similar to the performance with retrospective data. The average value of the area under the receiver operating characteristic curve (AUROC) decreased slightly by 0.6 percentage points (from 94.8% to 94.2% at discharge). The cross-hospital evaluations exhibited severely reduced performance: the average AUROC decreased by 8 percentage points (from 94.2% to 86.3% at discharge), which indicates the importance of model calibration with data from the deployment hospital. Conclusions Calibrating the prediction model with data from different deployment hospitals led to good performance in live settings. The performance degradation in the cross-hospital evaluation identified limitations in developing a generic model for different hospitals. Designing a generic process for model development to generate specialized prediction models for each hospital guarantees model performance in different hospitals.

Publisher

JMIR Publications Inc.

Subject

Health Informatics

Reference34 articles.

1. A guide to deep learning in healthcare

2. Machine intelligence in healthcare—perspectives on trustworthiness, explainability, usability, and transparency

3. Machine Learning in Medicine

4. Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic review

5. High-performance medicine: the convergence of human and artificial intelligence

Cited by 16 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Evaluating gender bias in ML-based clinical risk prediction models: A study on multiple use cases at different hospitals;Journal of Biomedical Informatics;2024-09

2. Assessing calibration and bias of a deployed machine learning malnutrition prediction model within a large healthcare system;npj Digital Medicine;2024-06-06

3. Machine Learning for Predicting Risk and Prognosis of Acute Kidney Disease in Critically Ill Elderly Patients During Hospitalization: Internet-Based and Interpretable Model Study;J MED INTERNET RES;2024

4. Machine Learning for Predicting Risk and Prognosis of Acute Kidney Disease in Critically Ill Elderly Patients During Hospitalization: Internet-Based and Interpretable Model Study;Journal of Medical Internet Research;2024-05-01

5. Advancing Precision Medicine: A Review of Innovative In Silico Approaches for Drug Development, Clinical Pharmacology and Personalized Healthcare;Pharmaceutics;2024-02-27