Systematic Review of Approaches to Preserve Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine-Reference-Cited by-同舟云学术

Systematic Review of Approaches to Preserve Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine

Published:2021-08 Issue:04 Volume:12 Page:808-815
ISSN:1869-0327
Container-title:Applied Clinical Informatics
language:en
Short-container-title:Appl Clin Inform

Author:

Guo Lin Lawrence¹,Pfohl Stephen R.²,Fries Jason²,Posada Jose²,Fleming Scott Lanyon²,Aftandilian Catherine³,Shah Nigam²,Sung Lillian¹⁴

Affiliation:

1. Program in Child Health Evaluative Sciences, The Hospital for Sick Children, Toronto, Canada

2. Biomedical Informatics Research, Stanford University, Palo Alto, California, United States

3. Division of Pediatric Hematology/Oncology, Stanford University, Palo Alto, United States

4. Division of Haematology/Oncology, The Hospital for Sick Children, Toronto, Canada

Abstract

Abstract Objective The change in performance of machine learning models over time as a result of temporal dataset shift is a barrier to machine learning-derived models facilitating decision-making in clinical practice. Our aim was to describe technical procedures used to preserve the performance of machine learning models in the presence of temporal dataset shifts. Methods Studies were included if they were fully published articles that used machine learning and implemented a procedure to mitigate the effects of temporal dataset shift in a clinical setting. We described how dataset shift was measured, the procedures used to preserve model performance, and their effects. Results Of 4,457 potentially relevant publications identified, 15 were included. The impact of temporal dataset shift was primarily quantified using changes, usually deterioration, in calibration or discrimination. Calibration deterioration was more common (n = 11) than discrimination deterioration (n = 3). Mitigation strategies were categorized as model level or feature level. Model-level approaches (n = 15) were more common than feature-level approaches (n = 2), with the most common approaches being model refitting (n = 12), probability calibration (n = 7), model updating (n = 6), and model selection (n = 6). In general, all mitigation strategies were successful at preserving calibration but not uniformly successful in preserving discrimination. Conclusion There was limited research in preserving the performance of machine learning models in the presence of temporal dataset shift in clinical medicine. Future research could focus on the impact of dataset shift on clinical decision making, benchmark the mitigation strategies on a wider range of datasets and tasks, and identify optimal strategies for specific settings.

Publisher

Georg Thieme Verlag KG

Subject

Health Information Management,Computer Science Applications,Health Informatics

Link

http://www.thieme-connect.de/products/ejournals/pdf/10.1055/s-0041-1735184.pdf

Reference29 articles.

1. The proliferation of reports on clinical scoring systems: issues about uptake and clinical utility;D W Challener;JAMA,2019

2. Scalable and accurate deep learning with electronic health records;A Rajkomar;NPJ Digit Med,2018

3. Multitask learning and benchmarking with clinical time series data;H Harutyunyan;Sci Data,2019

4. Barriers to Achieving Economies of Scale in Analysis of EHR Data. A Cautionary Tale;M P Sendak;Appl Clin Inform,2017

5. Machine intelligence in healthcare-perspectives on trustworthiness, explainability, usability, and transparency;C M Cutillo;NPJ Digit Med,2020

Cited by 34 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Performance of risk models to predict mortality risk for patients with heart failure: evaluation in an integrated health system;Clinical Research in Cardiology;2024-04-02

2. Sustainable deployment of clinical prediction tools—a 360° approach to model maintenance;Journal of the American Medical Informatics Association;2024-02-29

3. Deep continual multitask out-of-hospital incident severity assessment from changing clinical features;2024-02-22

4. Characterizing the limitations of using diagnosis codes in the context of machine learning for healthcare;BMC Medical Informatics and Decision Making;2024-02-14

5. Monitoring performance of clinical artificial intelligence: a scoping review protocol;JBI Evidence Synthesis;2024-02-08