wCM based hybrid pre-processing algorithm for class imbalanced dataset-Reference-Cited by-同舟云学术

wCM based hybrid pre-processing algorithm for class imbalanced dataset

Published:2021-09-15 Issue:2 Volume:41 Page:3339-3354
ISSN:1064-1246
Container-title:Journal of Intelligent & Fuzzy Systems
language:
Short-container-title:IFS

Author:

Singh Deepika¹,Saha Anju¹,Gosain Anjana¹

Affiliation:

1. USICT, Guru Gobind Singh Indraprasth University, Sector-16, C, Dwarka, New Delhi, India

Abstract

Imbalanced dataset classification is challenging because of the severely skewed class distribution. The traditional machine learning algorithms show degraded performance for these skewed datasets. However, there are additional characteristics of a classification dataset that are not only challenging for the traditional machine learning algorithms but also increase the difficulty when constructing a model for imbalanced datasets. Data complexity metrics identify these intrinsic characteristics, which cause substantial deterioration of the learning algorithms’ performance. Though many research efforts have been made to deal with class noise, none of them focused on imbalanced datasets coupled with other intrinsic factors. This paper presents a novel hybrid pre-processing algorithm focusing on treating the class-label noise in the imbalanced dataset, which suffers from other intrinsic factors such as class overlapping, non-linear class boundaries, small disjuncts, and borderline examples. This algorithm uses the wCM complexity metric (proposed for imbalanced dataset) to identify noisy, borderline, and other difficult instances of the dataset and then intelligently handles these instances. Experiments on synthetic datasets and real-world datasets with different levels of imbalance, noise, small disjuncts, class overlapping, and borderline examples are conducted to check the effectiveness of the proposed algorithm. The experimental results show that the proposed algorithm offers an interesting alternative to popular state-of-the-art pre-processing algorithms for effectively handling imbalanced datasets along with noise and other difficulties.

Publisher

IOS Press

Subject

Artificial Intelligence,General Engineering,Statistics and Probability

Reference36 articles.

1. A survey of predictive modeling on imbalanced domains;Branco;ACM Comput. Surv.,2016

2. A survey of multiple classifier systems as hybrid systems;Wozniak;Information Fusion,2014

3. Extreme entropy machines: robust information theoretic classification;Czarnecki;Pattern Anal. Appl.,2017

4. Paired feature multilayer ensemble- concept and evaluation of a classifier;Ksieniewicz;J. Intelligent and Fuzzy Systems,2017

5. Hybrid Data-Level Techniques for Class Imbalance Problem

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Comprehensive empirical investigation for prioritizing the pipeline of using feature selection and data resampling techniques;Journal of Intelligent & Fuzzy Systems;2024-03-05