Solving the class imbalance problem using ensemble algorithm: application of screening for aortic dissection-Reference-Cited by-同舟云学术

Solving the class imbalance problem using ensemble algorithm: application of screening for aortic dissection

Published:2022-03-28 Issue:1 Volume:22 Page:
ISSN:1472-6947
Container-title:BMC Medical Informatics and Decision Making
language:en
Short-container-title:BMC Med Inform Decis Mak

Author:

Liu Lijue,Wu Xiaoyu,Li Shihao,Li Yi,Tan Shiyang,Bai Yongping

Abstract

Abstract Background Imbalance between positive and negative outcomes, a so-called class imbalance, is a problem generally found in medical data. Despite various studies, class imbalance has always been a difficult issue. The main objective of this study was to find an effective integrated approach to address the problems posed by class imbalance and to validate the method in an early screening model for a rare cardiovascular disease aortic dissection (AD). Methods Different data-level methods, cost-sensitive learning, and the bagging method were combined to solve the problem of low sensitivity caused by the imbalance of two classes of data. First, feature selection was applied to select the most relevant features using statistical analysis, including significance test and logistic regression. Then, we assigned two different misclassification cost values for two classes, constructed weak classifiers based on the support vector machine (SVM) model, and integrated the weak classifiers with undersampling and bagging methods to build the final strong classifier. Due to the rarity of AD, the data imbalance was particularly prominent. Therefore, we applied our method to the construction of an early screening model for AD disease. Clinical data of 523,213 patients from the Institute of Hypertension, Xiangya Hospital, Central South University were used to verify the validity of this method. In these data, the sample ratio of AD patients to non-AD patients was 1:65, and each sample contained 71 features. Results The proposed ensemble model achieved the highest sensitivity of 82.8%, with training time and specificity reaching 56.4 s and 71.9% respectively. Additionally, it obtained a small variance of sensitivity of 19.58 × 10–3 in the seven-fold cross validation experiment. The results outperformed the common ensemble algorithms of AdaBoost, EasyEnsemble, and Random Forest (RF) as well as the single machine learning (ML) methods of logistic regression, decision tree, k nearest neighbors (KNN), back propagation neural network (BP) and SVM. Among the five single ML algorithms, the SVM model after cost-sensitive learning method performed best with a sensitivity of 79.5% and a specificity of 73.4%. Conclusions In this study, we demonstrate that the integration of feature selection, undersampling, cost-sensitive learning and bagging methods can overcome the challenge of class imbalance in a medical dataset and develop a practical screening model for AD, which could lead to a decision support for screening for AD at an early stage.

Publisher

Springer Science and Business Media LLC

Subject

Health Informatics,Health Policy,Computer Science Applications

Link

https://link.springer.com/content/pdf/10.1186/s12911-022-01821-w.pdf

Reference55 articles.

1. Belarouci S, Chikh MA. Medical imbalanced data classification. Adv Sci Technol Eng Syst J. 2017;2(3):116–24.

2. Bi J, Zhang C. An empirical comparison on state-of-the-art multi-class imbalance learning algorithms and a new diversified ensemble learning scheme. Knowl Based Syst. 2018;158(15):81–93.

3. Wu J, Zhao Z, Sun C, Yan R, Chen X. Learning from class-imbalanced data with a model-agnostic framework for machine intelligent diagnosis. Reliab Eng Syst Saf. 2021:107934.

4. Liu X-Y. An empirical study of boosting methods on severely imbalanced data. In: International conference on advances in materials science and information technologies in industry (AMSITI); 2014; Xian, Peoples R China.

5. Liu XY, Wu J, Zhou ZH. Exploratory undersampling for class-imbalance learning. IEEE Trans Syst Man Cybern. 2009;39(2):539–50.

Cited by 23 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Interpretable Support Vector Machine and Its Application to Rehabilitation Assessment;Electronics;2024-09-10

2. Borderline-DEMNET: A Workflow for Detecting Alzheimer’s and Dementia Stage by Solving Class Imbalance Problem;Pertanika Journal of Science and Technology;2024-07-16

3. Borderline-DEMNET: A Workflow for Detecting Alzheimer’s and Dementia Stage by Solving Class Imbalance Problem;Pertanika Journal of Science and Technology;2024-07-16

4. E-SMOTE: Entropy Based Minority Oversampling for Heart Failure and AIDS Clinical Trails Analysis;2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC);2024-07-02

5. Exploratory risk prediction of type II diabetes with isolation forests and novel biomarkers;Scientific Reports;2024-06-22