Validating the Early Prediction Model for COPD Patients Care through a Federated Machine Learning Architecture on FAIR Data (Preprint)-Reference-Cited by-同舟云学术

Validating the Early Prediction Model for COPD Patients Care through a Federated Machine Learning Architecture on FAIR Data (Preprint)

Published:2021-11-30 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

ALVAREZ-ROMERO Celia,MARTÍNEZ-GARCÍA Alicia,TERNERO-VEGA Jara Eloisa,DÍAZ-JIMÉNEZ Pablo,JIMÉNEZ-DE-JUAN Carlos,NIETO-MARTÍN María Dolores,ROMÁN-VILLARÁN Esther,KOVACEVIC Tomi,BOKAN Darijo,HROMIS Sanja,DJEKIC MALBASA Jelena,BESLAC Suzana,ZARIC Bojan,GENCTURK Mert,SINACI A. Anil,OLLERO-BATURONE Manuel,PARRA-CALDERÓN Carlos Luis^ORCID

Abstract

BACKGROUND

Due to the nature of health data, its sharing and reuse for research are limited by legal, technical and ethical implications. In this sense, to address that challenge, and facilitate and promote the discovery of scientific knowledge, the FAIR (Findable, Accessible, Interoperable, and Reusable) principles help organizations to share research data in a secure, appropriate and useful way for other researchers.

OBJECTIVE

The objective of this study was the FAIRification of health research existing datasets and applying a federated machine learning architecture on top of the FAIRified datasets of different health research performing organizations. The whole FAIR4Health solution was validated through the assessment of the generated model for real-time prediction of 30-days readmission risk in patients with Chronic Obstructive Pulmonary Disease (COPD).

METHODS

The application of the FAIR principles in health research datasets in three different health care settings enabled a retrospective multicenter study for the generation of federated machine learning models, aiming to develop the early prediction model for 30-days readmission risk in COPD patients. This prediction model was implemented upon the FAIR4Health platform and, finally, an observational prospective study with 30-days follow-up was carried out in two health care centers from different countries. The same inclusion and exclusion criteria were used in both retrospective and prospective parts of the study.

RESULTS

The prediction model for the 30-days hospital readmission risk was trained using the retrospective data of 4.944 COPD patients. The assessment of the prediction model was performed using the data of 100 recruited (22 from Spain and 78 from Serbia) out of 2070 observed (records viewed) patients in total for the observational prospective study from April 2021 to September 2021. The significant accuracy (0.98) and precision (0.25) of the prediction model generated upon the FAIR4Health platform was observed and, as a result, the generated prediction of 30-day readmission risk was confirmed in 87% of the cases.

CONCLUSIONS

A clinical validation was demonstrated through the implementation of federated machine learning models on top of the FAIRified datasets from different health research performing organizations, providing an assessment for predicting 30-days readmission risk in COPD patients. This demonstration allowed to state the relevance and need of implementing a FAIR data policy to facilitate data sharing and reuse in health research.

Publisher

JMIR Publications Inc.

Reference34 articles.

1. The FAIR Guiding Principles for scientific data management and stewardship

2. The Challenge of the Effective Implementation of FAIR Principles in Biomedical Research

3. Security and Privacy when Applying FAIR Principles to Genomic Information

4. A beginner’s guide to data stewardship and data sharing

5. A funder-imposed data publication requirement seldom inspired data sharing