Generating Synthetic Electronic Health Record Data Using Generative Adversarial Networks: Tutorial-Reference-Cited by-同舟云学术

Generating Synthetic Electronic Health Record Data Using Generative Adversarial Networks: Tutorial

Published:2024-04-22 Issue: Volume:3 Page:e52615
ISSN:2817-1705
Container-title:JMIR AI
language:en
Short-container-title:JMIR AI

Author:

Yan Chao^ORCID,Zhang Ziqi^ORCID,Nyemba Steve^ORCID,Li Zhuohang^ORCID

Abstract

Synthetic electronic health record (EHR) data generation has been increasingly recognized as an important solution to expand the accessibility and maximize the value of private health data on a large scale. Recent advances in machine learning have facilitated more accurate modeling for complex and high-dimensional data, thereby greatly enhancing the data quality of synthetic EHR data. Among various approaches, generative adversarial networks (GANs) have become the main technical path in the literature due to their ability to capture the statistical characteristics of real data. However, there is a scarcity of detailed guidance within the domain regarding the development procedures of synthetic EHR data. The objective of this tutorial is to present a transparent and reproducible process for generating structured synthetic EHR data using a publicly accessible EHR data set as an example. We cover the topics of GAN architecture, EHR data types and representation, data preprocessing, GAN training, synthetic data generation and postprocessing, and data quality evaluation. We conclude this tutorial by discussing multiple important issues and future opportunities in this domain. The source code of the entire process has been made publicly available.

Publisher

JMIR Publications Inc.

Reference43 articles.

1. Synthetic patient data in health care: a widening legal loophole

2. Synthetic data in health care: A narrative review

3. Synthetic data in machine learning for medicine and healthcare

4. Big Data and Machine Learning in Health Care

5. Big Data in Public Health: Terminology, Machine Learning, and Privacy