Synthetic Health Data Can Augment Community Research Efforts to Better Inform the Public During Emerging Pandemics-Reference-Cited by-同舟云学术

Synthetic Health Data Can Augment Community Research Efforts to Better Inform the Public During Emerging Pandemics

Published:2023-12-13 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Prasanna Anish,Jing Bocheng,Plopper George,Miller Kristina Krasnov,Sanjak Jaleal,Feng Alice,Prezek Sarah,Vidyaprakash Eshaw,Thovarai Vishal,Maier Ezekiel J.^ORCID,Bhattacharya Avik,Naaman Lama,Stephens Holly,Watford Sean,John Boscardin W.,Johanson Elaine,Lienau Amanda

Abstract

ABSTRACTThe COVID-19 pandemic had disproportionate effects on the Veteran population due to the increased prevalence of medical and environmental risk factors. Synthetic electronic health record (EHR) data can help meet the acute need for Veteran population-specific predictive modeling efforts by avoiding the strict barriers to access, currently present within Veteran Health Administration (VHA) datasets. The U.S. Food and Drug Administration (FDA) and the VHA launched the precisionFDA COVID-19 Risk Factor Modeling Challenge to develop COVID-19 diagnostic and prognostic models; identify Veteran population-specific risk factors; and test the usefulness of synthetic data as a substitute for real data. The use of synthetic data boosted challenge participation by providing a dataset that was accessible to all competitors. Models trained on synthetic data showed similar but systematically inflated model performance metrics to those trained on real data. The important risk factors identified in the synthetic data largely overlapped with those identified from the real data, and both sets of risk factors were validated in the literature. Tradeoffs exist between synthetic data generation approaches based on whether a real EHR dataset is required as input. Synthetic data generated directly from real EHR input will more closely align with the characteristics of the relevant cohort. This work shows that synthetic EHR data will have practical value to the Veterans’ health research community for the foreseeable future.

Publisher

Cold Spring Harbor Laboratory

Reference29 articles.

1. Beyond the hype of big data and artificial intelligence: building foundations for knowledge and wisdom;BMC Med,2019

2. Explainable artificial intelligence models using real-world electronic health record data: a systematic scoping review

3. Real-world data mining meets clinical practice: Research challenges and perspective;Front Big Data,2022

4. Leveraging electronic health records for data science: common pitfalls and how to avoid them;The Lancet Digital Health,2022

5. From real-world electronic health record data to real-world results using artificial intelligence

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A framework for sharing of clinical and genetic data for precision medicine applications;Nature Medicine;2024-09-03