A self-inspected adaptive SMOTE algorithm (SASMOTE) for highly imbalanced data classification in healthcare-Reference-Cited by-同舟云学术

A self-inspected adaptive SMOTE algorithm (SASMOTE) for highly imbalanced data classification in healthcare

Published:2023-04-25 Issue:1 Volume:16 Page:
ISSN:1756-0381
Container-title:BioData Mining
language:en
Short-container-title:BioData Mining

Author:

Kosolwattana Tanapol,Liu Chenang,Hu Renjie,Han Shizhong,Chen Hua,Lin Ying

Abstract

AbstractIn many healthcare applications, datasets for classification may be highly imbalanced due to the rare occurrence of target events such as disease onset. The SMOTE (Synthetic Minority Over-sampling Technique) algorithm has been developed as an effective resampling method for imbalanced data classification by oversampling samples from the minority class. However, samples generated by SMOTE may be ambiguous, low-quality and non-separable with the majority class. To enhance the quality of generated samples, we proposed a novel self-inspected adaptive SMOTE (SASMOTE) model that leverages an adaptive nearest neighborhood selection algorithm to identify the “visible” nearest neighbors, which are used to generate samples likely to fall into the minority class. To further enhance the quality of the generated samples, an uncertainty elimination via self-inspection approach is introduced in the proposed SASMOTE model. Its objective is to filter out the generated samples that are highly uncertain and inseparable with the majority class. The effectiveness of the proposed algorithm is compared with existing SMOTE-based algorithms and demonstrated through two real-world case studies in healthcare, including risk gene discovery and fatal congenital heart disease prediction. By generating the higher quality synthetic samples, the proposed algorithm is able to help achieve better prediction performance (in terms of F1 score) on average compared to the other methods, which is promising to enhance the usability of machine learning models on highly imbalanced healthcare data.

Funder

National Institute of Mental Health

Publisher

Springer Science and Business Media LLC

Subject

Computational Mathematics,Computational Theory and Mathematics,Computer Science Applications,Genetics,Molecular Biology,Biochemistry

Link

https://link.springer.com/content/pdf/10.1186/s13040-023-00330-4.pdf

Reference31 articles.

1. Sun Y, Wong AK, Kamel MS. Classification of imbalanced data: A review. Int J Pattern Recognit Artif Intell. 2009;23(04):687–719.

2. Zhao Y, Wong ZSY, Tsui KL. A framework of rebalancing imbalanced healthcare data for rare events’ classification: a case of look-alike sound-alike mix-up incident detection. J Healthc Eng. 2018;2018:1–11. https://doi.org/10.1155/2018/6275435.

3. Nakamura M, Kajiwara Y, Otsuka A, Kimura H. Lvq-smote-learning vector quantization based synthetic minority over-sampling technique for biomedical data. BioData Min. 2013;6(1):1–10.