Enabling qualitative research data sharing using a natural language processing pipeline for deidentification: moving beyond HIPAA Safe Harbor identifiers-Reference-Cited by-同舟云学术

Enabling qualitative research data sharing using a natural language processing pipeline for deidentification: moving beyond HIPAA Safe Harbor identifiers

Published:2021-07-01 Issue:3 Volume:4 Page:
ISSN:2574-2531
Container-title:JAMIA Open
language:en
Short-container-title:

Author:

Gupta Aditi¹,Lai Albert¹^ORCID,Mozersky Jessica²,Ma Xiaoteng¹,Walsh Heidi²,DuBois James M²^ORCID

Affiliation:

1. Institute for Informatics, Washington University, St. Louis, Missouri, USA

2. Bioethics Research Center, Division of General Medical Sciences, Washington University, St. Louis, Missouri, USA

Abstract

Abstract Objective Sharing health research data is essential for accelerating the translation of research into actionable knowledge that can impact health care services and outcomes. Qualitative health research data are rarely shared due to the challenge of deidentifying text and the potential risks of participant reidentification. Here, we establish and evaluate a framework for deidentifying qualitative research data using automated computational techniques including removal of identifiers that are not considered HIPAA Safe Harbor (HSH) identifiers but are likely to be found in unstructured qualitative data. Materials and Methods We developed and validated a pipeline for deidentifying qualitative research data using automated computational techniques. An in-depth analysis and qualitative review of different types of qualitative health research data were conducted to inform and evaluate the development of a natural language processing (NLP) pipeline using named-entity recognition, pattern matching, dictionary, and regular expression methods to deidentify qualitative texts. Results We collected 2 datasets with 1.2 million words derived from over 400 qualitative research data documents. We created a gold-standard dataset with 280K words (70 files) to evaluate our deidentification pipeline. The majority of identifiers in qualitative data are non-HSH and not captured by existing systems. Our NLP deidentification pipeline had a consistent F1-score of ∼0.90 for both datasets. Conclusion The results of this study demonstrate that NLP methods can be used to identify both HSH identifiers and non-HSH identifiers. Automated tools to assist researchers with the deidentification of qualitative data will be increasingly important given the new National Institutes of Health (NIH) data-sharing mandate.

Funder

National Human Genome Research Institute of the U.S. National Institutes of Health

National Center for Advancing Translational Sciences

National Institutes of Health or the National Human Genome Research Institute

Publisher

Oxford University Press (OUP)

Subject

Health Informatics

Link

http://academic.oup.com/jamiaopen/article-pdf/4/3/ooab069/39868051/ooab069.pdf

Reference38 articles.

1. The role of qualitative research in HIV/AIDS;Power;AIDS,1998

2. Qualitative research and its uses in health care;Al-Busaidi;Sultan Qaboos Univ Med J,2008

3. Are we ready to share qualitative research data? Knowledge and preparedness among qualitative researchers, IRB members, and data repository curators;Mozersky;IASSIST Q,2020

4. Is it time to share qualitative research data?;DuBois;Qual Psychol,2018

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Natural Language Processing in Healthcare;Artificial and Cognitive Computing for Sustainable Healthcare Systems in Smart Cities;2024-05-17

2. Exchanging words: Engaging the challenges of sharing qualitative research data;Proceedings of the National Academy of Sciences;2023-10-13

3. Open-Science Guidance for Qualitative Research: An Empirically Validated Approach for De-Identifying Sensitive Narrative Data;Advances in Methods and Practices in Psychological Science;2023-10

4. How might responsible management education (RME) be used to develop responsible leadership skills among students in business schools? Evidence from non-Western business schools;European Journal of Training and Development;2023-09-01

5. ChatGPT: Can a Natural Language Processing Tool Be Trusted for Radiation Oncology Use?;International Journal of Radiation Oncology*Biology*Physics;2023-08