Affiliation:
1. Department of Computer Science and Engineering, Kumaraguru College of Technology, Coimbatore, Tamil Nadu, India
2. Department of Electrical and Computer Engineering, Faculty of Engineering, King Abdulaziz University, Jeddah, Kingdom of Saudi Arabia
Abstract
Big Data is a popular research area where a vast amount of data is created, replicated, and consumed by society. The quality of the data used directly influences big data knowledge discovery. The existence of noise is the most prevalent problem influencing data quality. The following techniques were developed to reduce noise in data with a distributed setting: Homogenous Ensemble for Big Data (HME-BD) and Heterogeneous Ensemble for Big Data (HTE-BD). In this article, the performance of HTE-BD is improved further by developing Enhanced HTE-BD (EHTE-BD), which combines Logistic Regression based Support Vector Machine (LR-SVM) in conjunction with RF, LR, and KNN to reduce noisy data. Furthermore, the Multi-Objective Evolutionary Fuzzy Method for Subgroup Discovery throughout Big Data (MEFASD-BD) was used to resolve the multi-objective optimization challenge, and the Non-Dominated Sorting Genetic Algorithm-II (NSGA-II) was utilized to handle the rising dimensionality issue through subgroup discovery. To address the NSGA-II’s slow convergence rate, an Improved Multi-Objective Meta-Heuristic Fuzzy approach for discovering subgroups in big data is described, that contains a meta-heuristic method for subgroup discovery known as the Multi-Objective Differential Search Algorithm (MODSA). It selects the most relevant subgroups from vast amounts of data, reducing the data’s dimensionality. The Fuzzy Deep Neural Network (FDNN) classifier assesses the main subgroups. By removing noisy data and selecting the most relevant subgroups, the performance of FDNN in classifying vast amounts of data is improved.
Subject
Artificial Intelligence,General Engineering,Statistics and Probability
Reference38 articles.
1. Data mining with big data;Wu;IEEE Transactions on Knowledge and Data Engineering,2014
2. Data-intensive applications, challenges, techniques, and technologies: A survey on Big Data;Chen;Information Sciences,2014
3. Business intelligence and analytics: from big data to big impact;Chen;M.I.S.,2012
4. Big data: tutorial andguidelines on information and process fusion for analyticsalgorithms with Map Reduce;Ramírez-Gallego;Inform Fusion,2018
5. Big data with cloud computing: an insight on the computing environment, map-reduce, and programming frameworks;Fernandez;WIRES: Data Min Know Discov,2014