Abstract
Addressing the class imbalance in classification problems is particularly challenging, especially in the context of medical datasets where misclassifying minority class samples can have significant repercussions. This study is dedicated to mitigating class imbalance in medical datasets by employing a hybrid approach that combines data-level, cost-sensitive, and ensemble methods. Through an assessment of the performance, measured by AUC-ROC values, Sensitivity, F1-Score, and G-Mean of 20 data-level and four cost-sensitive models on seventeen medical datasets - 12 small and five large, a hybridized model, SMOTE-RF-CS-LR has been devised. This model integrates the Synthetic Minority Oversampling Technique (SMOTE), the ensemble classifier Random Forest (RF), and the Cost-Sensitive Logistic Regression (CS-LR). Upon testing the hybridized model on diverse imbalanced ratios, it demonstrated remarkable performance, achieving outstanding performance values on the majority of the datasets. Further examination of the model's training duration and time complexity revealed its efficiency, taking less than a second to train on each small dataset. Consequently, the proposed hybridized model not only proves to be time-efficient but also exhibits robust capabilities in handling class imbalance, yielding outstanding classification results in the context of medical datasets.
Publisher
Asian Research Association
Reference45 articles.
1. N. Japkowicz, S. Stephen, The class imbalance problem: A systematic study. Intelligent data analysis, 6(5), (2002) 429-449.
2. A. Ali, S.M. Shamsuddin, A.L. Ralescu, Classification with class imbalance problem. International Journal of Advances in Soft Computing and its Applications, 5(3), (2013) 176–204.
3. M.C. Monard, G. Batista, Learning with skewed class distributions. Advances in Logic, Artificial Intelligence and Robotics, 85, (2002) 173–180.
4. J. Tanha, Y. Abdi, N. Samadi, N. Razzaghi, M. Asadpour, Boosting methods for multi-class imbalanced data classification: an experimental review. Journal of Big Data, 7, (2020) 1–47. https://doi.org/10.1186/s40537-020-00349-y
5. G. Aguiar, B. Krawczyk, A. Cano, A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework. Machine Learning, (2023) 1–79. https://doi.org/10.1007/s10994-023-06353-6