Benchmarking Machine Learning Missing Data Imputation Methods in Large-Scale Mental Health Survey Databases-Reference-Cited by-同舟云学术

Benchmarking Machine Learning Missing Data Imputation Methods in Large-Scale Mental Health Survey Databases

Published:2024-05-14 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Prakash Preethi,Street Kelly,Narayanan Shrikanth,Fernandez Bridget A.,Shen Yufeng^ORCID,Shu Chang^ORCID

Abstract

AbstractDatabases with mental and behavioral health surveys suffer from missingness when participants skip the entire survey, affecting the data quality and sample size. We investigated the missing data patterns and evaluate the imputation performance in Simons Powering Autism Research (SPARK), a large-scale autism cohort consists of over 117,000 participants. Four common methods were assessed – Multiple Imputation by Chained Equations (MICE), K-Nearest Neighbors (KNN), MissForest, and Multiple Imputation with Denoising Autoencoders (MIDAS). In a complete subset of 15,196 autism participants, we simulated three types of missingness patterns. We observed that MIDAS and KNN performed the best as the rate of random missingness increased and when blockwise missingness was simulated. The average computational times for MIDAS and KNN were 10 minutes, 35 minutes for MissForest, and 290 minutes for MICE. MIDAS and KNN both provide promising imputation performance in mental and behavioral health survey data that exhibit blockwise missingness patterns.

Publisher

Cold Spring Harbor Laboratory

Reference27 articles.

1. SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research

2. Mental health in UK Biobank – development, implementation and results from an online questionnaire completed by 157 366 participants: a reanalysis

3. The All of Us Research Program: Data quality, utility, and diversity

4. A meta-analysis of the social communication questionnaire: Screening for autism spectrum disorder

5. Psychometric analysis of the repetitive behavior scale‐revised using confirmatory factor analysis in children with autism