Missing data imputation, prediction, and feature selection in diagnosis of vaginal prolapse-Reference-Cited by-同舟云学术

Missing data imputation, prediction, and feature selection in diagnosis of vaginal prolapse

Published:2023-11-06 Issue:1 Volume:23 Page:
ISSN:1471-2288
Container-title:BMC Medical Research Methodology
language:en
Short-container-title:BMC Med Res Methodol

Author:

FAN Mingxuan,Peng Xiaoling,Niu Xiaoyu,Cui Tao,He Qiaolin

Abstract

Abstract Background Data loss often occurs in the collection of clinical data. Directly discarding the incomplete sample may lead to low accuracy of medical diagnosis. A suitable data imputation method can help researchers make better use of valuable medical data. Methods In this paper, five popular imputation methods including mean imputation, expectation-maximization (EM) imputation, K-nearest neighbors (KNN) imputation, denoising autoencoders (DAE) and generative adversarial imputation nets (GAIN) are employed on an incomplete clinical data with 28,274 cases for vaginal prolapse prediction. A comprehensive comparison study for the performance of these methods has been conducted through certain classification criteria. It is shown that the prediction accuracy can be greatly improved by using the imputed data, especially by GAIN. To find out the important risk factors to this disease among a large number of candidate features, three variable selection methods: the least absolute shrinkage and selection operator (LASSO), the smoothly clipped absolute deviation (SCAD) and the broken adaptive ridge (BAR) are implemented in logistic regression for feature selection on the imputed datasets. In pursuit of our primary objective, which is accurate diagnosis, we employed diagnostic accuracy (classification accuracy) as a pivotal metric to assess both imputation and feature selection techniques. This assessment encompassed seven classifiers (logistic regression (LR) classifier, random forest (RF) classifier, support machine classifier (SVC), extreme gradient boosting (XGBoost) , LASSO classifier, SCAD classifier and Elastic Net classifier)enhancing the comprehensiveness of our evaluation. Results The proposed framework imputation-variable selection-prediction is quite suitable to the collected vaginal prolapse datasets. It is observed that the original dataset is well imputed by GAIN first, and then 9 most significant features were selected using BAR from the original 67 features in GAIN imputed dataset, with only negligible loss in model prediction. BAR is superior to the other two variable selection methods in our tests. Concludes Overall, combining the imputation, classification and variable selection, we achieve good interpretability while maintaining high accuracy in computer-aided medical diagnosis.

Funder

National Key R \& D Program of China

Science \& Technology of Sichuan

Publisher

Springer Science and Business Media LLC

Subject

Health Informatics,Epidemiology

Link

https://link.springer.com/content/pdf/10.1186/s12874-023-02079-0.pdf

Reference51 articles.

1. Jelovsek JE, Maher C, Barber MD. Pelvic organ prolapse. Lancet. 2007;369(9566):1027–38.

2. Pang H, Zhang L, Han S, Li Z, Gong J, Liu Q, et al. A nationwide population-based survey on the prevalence and risk factors of symptomatic pelvic organ prolapse in adult women in China-a pelvic organ prolapse quantification system-based study. BJOG Int J Obstet Gynaecol. 2021;128(8):1313–23.

3. Olsen AL, Smith VJ, Bergstrom JO, Colling JC, Clark AL. Epidemiology of surgically managed pelvic organ prolapse and urinary incontinence. Obstet Gynecol. 1997;89(4):501–6.

4. Jerez JM, Molina I, García-Laencina PJ, Alba E, Ribelles N, Martín M, et al. Missing data imputation using statistical and machine learning methods in a real breast cancer problem. Artif Intell Med. 2010;50(2):105–15.

5. Nagarajan G, Babu LD. Missing data imputation on biomedical data using deeply learned clustering and L2 regularized regression based on symmetric uncertainty. Artif Intell Med. 2022;123:102214.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Handling missing data and measurement error for early-onset myopia risk prediction models;BMC Medical Research Methodology;2024-09-06