EKMGS: A HYBRID CLASS BALANCING METHOD FOR MEDICAL DATA PROCESSING-Reference-Cited by-同舟云学术

EKMGS: A HYBRID CLASS BALANCING METHOD FOR MEDICAL DATA PROCESSING

Published:2024-06-30 Issue: Volume: Page:5-16
ISSN:2707-904X
Container-title:Scientific Journal of Astana IT University
language:
Short-container-title:sjaitu

Author:

Buribayev Zholdas^ORCID,Shaikalamova Saida^ORCID,Yerkos Ainur^ORCID,Imanbek Rustem^ORCID

Abstract

The field of medicine is witnessing rapid development of AI, highlighting the importance of proper data processing. However, when working with medical data, there is a problem of class imbalance, where the amount of data about healthy patients significantly exceeds the amount of data about sick ones. This leads to incorrect classification of the minority class, resulting in inefficient operation of machine learning algorithms. In this study, a hybrid method was developed to address the problem of class imbalance, combining oversampling (GenSMOTE) and undersampling (ENN) algorithms. GenSMOTE used frequency oversampling optimization based on a genetic algorithm, selecting the optimal value using a fitness function. The next stage implemented an ensemble method based on stacking, consisting of three base (k-NN, SVM, LR) and one meta-model (Decision Tree). The hyperparameters of the meta-model were optimized using the GridSearchCV algorithm. During the study, datasets on diabetes, liver diseases, and brain glioma were used. The developed hybrid class balancing method significantly improved the quality of the model: the F1-score increased by 10-75%, and accuracy by 5-30%. Each stage of the hybrid algorithm was visualized using a nonlinear UMAP algorithm. The ensemble method based on stacking, in combination with the hybrid class balancing method, demonstrated high efficiency in solving classification tasks in medicine. This approach can be applied for diagnosing various diseases, which will increase the accuracy and reliability of forecasts. It is planned to expand the application of this approach to large volumes of data and improve the oversampling algorithm using additional capabilities of the genetic algorithm.

Publisher

Astana IT University

Reference20 articles.

1. Xu, Z. , Shen, D. , Nie, T. , Kou, Y. , Yin, N. , & Han, X. (2021). A cluster-based oversampling algorithm combining SMOTE and k-means for imbalanced medical data. Information Sciences, 572, 574- 589. https://doi.org/10.1016/j.ins.2021.02.056

2. Khushi, M., Shaukat, K., Alam, T. M., Hameed, I. A., Uddin, S., Luo, S., Yang, X., & Reyes, M. C. (2021). A Comparative Performance Analysis of Data Resampling Methods on Imbalance Medical Data. IEEE Access, 9, 109960–109975.

3. Mienye, I.D., & Sun, Y. (2021). Performance analysis of cost-sensitive learning methods with ap- plication to imbalanced medical data. Informatics in Medicine Unlocked, 25, 100690. https://doi. org/10.1016/j.imu.2021.100690

4. Wang, Y.-C., & Cheng, C.-H. (2021). A multiple combined method for rebalancing medical data with class imbalances. Computers in Biology and Medicine, 134, 104527. https://doi.org/10.1016/j. compbiomed.2021.104527

5. Lee, D., & Kim, K. (2021). An efficient method to determine sample size in oversampling based on classification complexity for imbalanced data. Expert Systems with Applications, 184, 115442. https://doi.org/10.1016/j.eswa.2021.115442