Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies

Author:

Reiner Benaim AnatORCID,Almog RonitORCID,Gorelik YuriORCID,Hochberg IritORCID,Nassar LailaORCID,Mashiach TanyaORCID,Khamaisi MogherORCID,Lurie YaelORCID,Azzam Zaher SORCID,Khoury JohadORCID,Kurnik DanielORCID,Beyar RafaelORCID

Abstract

Background Privacy restrictions limit access to protected patient-derived health information for research purposes. Consequently, data anonymization is required to allow researchers data access for initial analysis before granting institutional review board approval. A system installed and activated at our institution enables synthetic data generation that mimics data from real electronic medical records, wherein only fictitious patients are listed. Objective This paper aimed to validate the results obtained when analyzing synthetic structured data for medical research. A comprehensive validation process concerning meaningful clinical questions and various types of data was conducted to assess the accuracy and precision of statistical estimates derived from synthetic patient data. Methods A cross-hospital project was conducted to validate results obtained from synthetic data produced for five contemporary studies on various topics. For each study, results derived from synthetic data were compared with those based on real data. In addition, repeatedly generated synthetic datasets were used to estimate the bias and stability of results obtained from synthetic data. Results This study demonstrated that results derived from synthetic data were predictive of results from real data. When the number of patients was large relative to the number of variables used, highly accurate and strongly consistent results were observed between synthetic and real data. For studies based on smaller populations that accounted for confounders and modifiers by multivariate models, predictions were of moderate accuracy, yet clear trends were correctly observed. Conclusions The use of synthetic structured data provides a close estimate to real data results and is thus a powerful tool in shaping research hypotheses and accessing estimated analyses, without risking patient privacy. Synthetic data enable broad access to data (eg, for out-of-organization researchers), and rapid, safe, and repeatable analysis of data in hospitals or other health organizations where patient privacy is a primary value.

Publisher

JMIR Publications Inc.

Subject

Health Information Management,Health Informatics

Reference37 articles.

1. GarfinkleSLNational Institute of Standards and Technology2015102020-01-20De-Identification of Personal Informationhttps://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf

2. GrahamCThe Information Commissioner's Office (ICO)20122020-01-20Anonymization: Managing Data Protection Risk Code of Practicehttps://ico.org.uk/media/for-organisations/documents/1061/anonymisation-code.pdf

3. Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record

4. Under threat: patient confidentiality and NHS computing

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3