Selective ensemble learning algorithm for imbalanced dataset-Reference-Cited by-同舟云学术

Selective ensemble learning algorithm for imbalanced dataset

Published:2023 Issue:2 Volume:20 Page:831-856
ISSN:1820-0214
Container-title:Computer Science and Information Systems
language:en
Short-container-title:COMSIS J

Author:

Du Hongle¹,Zhang Yan²,Zhang Lin²,Chen Yeh-Cheng³

Affiliation:

1. School of Mathematics and Computer Application, Shangluo University, Shangluo, China + University of the Cordilleras, aguio City, Philippines + Shangluo Public Big Data Research Center, Shangluo, China

2. School of Mathematics and Computer Application, Shangluo University, Shangluo, China + Shangluo Public Big Data Research Center, Shangluo, China

3. Department of computer science, University of California, Davis, USA

Abstract

Under the imbalanced dataset, the performance of the base-classifier, the computing method of weight of base-classifier and the selection method of the base-classifier have a great impact on the performance of the ensemble classifier. In order to solve above problem to improve the generalization performance of ensemble classifier, a selective ensemble learning algorithm based on under-sampling for imbalanced dataset is proposed. First, the proposed algorithm calculates the number K of under-sampling samples according to the relationship between class sample density. Then, we use the improved K-means clustering algorithm to under-sample the majority class samples and obtain K cluster centers. Then, all cluster centers (or the sample of the nearest cluster center) are regarded as new majority samples to construct a new balanced training subset combine with the minority class?s samples. Repeat those processes to generate multiple training subsets and get multiple base-classifiers. However, with the increasing of iterations, the number of base-classifiers increase, and the similarity among the base-classifiers will also increase. Therefore, it is necessary to select some base-classifier with good classification performance and large difference for ensemble. In the stage of selecting base-classifiers, according to the difference and performance of base-classifiers, we use the idea of maximum correlation and minimum redundancy to select base-classifiers. In the ensemble stage, G-mean or F-mean is selected to evaluate the classification performance of base-classifier for imbalanced dataset. That is to say, it is selected to compute the weight of each base-classifier. And then the weighted voting method is used for ensemble. Finally, the simulation results on the artificial dataset, UCI dataset and KDDCUP dataset show that the algorithm has good generalization performance on imbalanced dataset, especially on the dataset with high imbalance degree.

Publisher

National Library of Serbia

Subject

General Computer Science

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Multi-class imbalance problem: A multi-objective solution;Information Sciences;2024-10

2. Integrated optimization of line planning and timetabling on high-speed railway network considering cross-line operation;Advances in Production Engineering & Management;2024-03-29