Perceptions of Data Set Experts on Important Characteristics of Health Data Sets Ready for Machine Learning-Reference-Cited by-同舟云学术

Perceptions of Data Set Experts on Important Characteristics of Health Data Sets Ready for Machine Learning

Published:2023-12-01 Issue:12 Volume:6 Page:e2345892
ISSN:2574-3805
Container-title:JAMA Network Open
language:en
Short-container-title:JAMA Netw Open

Author:

Ng Madelena Y.¹²,Youssef Alaa³,Miner Adam S.⁴,Sarellano Daniela³,Long Jin⁵,Larson David B.³,Hernandez-Boussard Tina¹²,Langlotz Curtis P.¹²³

Affiliation:

1. Department of Medicine (Biomedical Informatics), Stanford University School of Medicine, Stanford, California

2. Department of Biomedical Data Science, Stanford University School of Medicine, Stanford, California

3. Department of Radiology, Stanford University School of Medicine, Stanford, California

4. Department of Psychiatry and Behavioral Sciences, Stanford University School of Medicine, Stanford, California

5. Department of Pediatrics, Stanford University School of Medicine, Stanford, California

Abstract

ImportanceThe lack of data quality frameworks to guide the development of artificial intelligence (AI)-ready data sets limits their usefulness for machine learning (ML) research in health care and hinders the diagnostic excellence of developed clinical AI applications for patient care.ObjectiveTo discern what constitutes high-quality and useful data sets for health and biomedical ML research purposes according to subject matter experts.Design, Setting, and ParticipantsThis qualitative study interviewed data set experts, particularly those who are creators and ML researchers. Semistructured interviews were conducted in English and remotely through a secure video conferencing platform between August 23, 2022, and January 5, 2023. A total of 93 experts were invited to participate. Twenty experts were enrolled and interviewed. Using purposive sampling, experts were affiliated with a diverse representation of 16 health data sets/databases across organizational sectors. Content analysis was used to evaluate survey information and thematic analysis was used to analyze interview data.Main Outcomes and MeasuresData set experts’ perceptions on what makes data sets AI ready.ResultsParticipants included 20 data set experts (11 [55%] men; mean [SD] age, 42 [11] years), of whom all were health data set creators, and 18 of the 20 were also ML researchers. Themes (3 main and 11 subthemes) were identified and integrated into an AI-readiness framework to show their association within the health data ecosystem. Participants partially determined the AI readiness of data sets using priority appraisal elements of accuracy, completeness, consistency, and fitness. Ethical acquisition and societal impact emerged as appraisal considerations in that participant samples have not been described to date in prior data quality frameworks. Factors that drive creation of high-quality health data sets and mitigate risks associated with data reuse in ML research were also relevant to AI readiness. The state of data availability, data quality standards, documentation, team science, and incentivization were associated with elements of AI readiness and the overall perception of data set usefulness.Conclusions and RelevanceIn this qualitative study of data set experts, participants contributed to the development of a grounded framework for AI data set quality. Data set AI readiness required the concerted appraisal of many elements and the balancing of transparency and ethical reflection against pragmatic constraints. The movement toward more reliable, relevant, and ethical AI and ML applications for patient care will inevitably require strategic updates to data set creation practices.

Publisher

American Medical Association (AMA)

Subject

General Medicine

Link

https://jamanetwork.com/journals/jamanetworkopen/articlepdf/2812417/ng_2023_oi_231335_1700596728.84438.pdf

Reference53 articles.

1. Machine learning in medicine.;Rajkomar;N Engl J Med,2019

2. High-performance medicine: the convergence of human and artificial intelligence.;Topol;Nat Med,2019

3. Clinical applications of artificial intelligence—an updated overview.;Busnatu;J Clin Med,2022

4. Ethics of using and sharing clinical imaging data for artificial intelligence: a proposed framework.;Larson;Radiology,2020

5. Transparency and reproducibility in artificial intelligence.;Haibe-Kains;Nature,2020

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Cross-modal hybrid architectures for gastrointestinal tract image analysis: A systematic review and futuristic applications;Image and Vision Computing;2024-08

2. Explainable artificial intelligence in breast cancer detection and risk prediction: A systematic scoping review;Cancer Innovation;2024-07-03

3. NNI nanoinformatics conference 2023: Movement toward a common infrastructure for federal nanoEHS data computational toxicology: Short communication;Computational Toxicology;2024-06

4. PROBAST Assessment of Machine Learning: Reply;Anesthesiology;2024-05-29

5. Machine learning for healthcare that matters: Reorienting from technical novelty to equitable impact;PLOS Digital Health;2024-04-15