Ensemble Model for Educational Data Mining Based on Synthetic Minority Oversampling Technique-Reference-Cited by-同舟云学术

Ensemble Model for Educational Data Mining Based on Synthetic Minority Oversampling Technique

Published:2023-06-14 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

R Manoharan¹^ORCID,Stalin M.Subi²,Loganathan Ganesh Babu³,s Deepa⁴,k Venkateswaran⁵

Affiliation:

1. Apollo Engineering College

2. Arignar Anna Institute of Science and Technology

3. Tishk International University

4. Kongu Engineering College

5. St Joseph Engineering College

Abstract

Abstract Data mining in the classroom is a well-known field that involves data mining concepts, statistical analysis, and machine learning concepts, all of which are applied to educational data. These EDM processed data are frequently used to analyse various aspects of the business and process model. Existing and traditional process models involve the use of traditional statistical techniques to process data, which necessitates a significant amount of manual intervention for data modelling and pre-processing. To address the issues raised above, this paper proposes a novel technique that combines a machine learning model with statistical approaches. This machine learning combination combines various classifiers such as Decision Tree Logistic Regression, Random Forest, Multiplayer Perceptron, K Nearest Neighbor, Decision Tree, and so on. Because of the limited data availability, the information used in the observation is highly imbalanced. As a result, the above-mentioned technique is combined with the universally benchmarked model known as Synthetic Minority Oversampling Technique (SMOTE) to discuss issues related to class imbalance. In addition, the performance evaluation is statistically performed to demonstrate the efficacy of the suggested strategies. A technological college in India obtained the main student data collection, which included information on 6,807 students with characteristics. A synthetic minority oversampling approach sensor is used to handle the imbalanced data set. The model is calibrated using eight methodologies, which are then evaluated to determine the dimensions that will help produce the best suitable model to categorise a student based on his achievements.

Publisher

Research Square Platform LLC

Reference24 articles.

1. Sapkota N, Alsadoon A, Prasad PWC, Elchouemi A, Singh AK (2019) "Data Summarization Using Clustering and Classification: Spectral Clustering Combined with k-Means Using NFPH," International Conference on Machine Learning, Big Data, Cloud and Parallel Computing (COMITCon), Faridabad, India, 2019, pp. 146–151, doi: 10.1109/COMITCon.2019.8862218

2. Li H, Lu Q (2017) "K-CV parameter optimization method in the application of SVM classification data," 2017 IEEE 2nd International Conference on Big Data Analysis (ICBDA), Beijing, pp. 25–29, doi: 10.1109/ICBDA.2017.8078838

3. Chandra S, Kaur M (2015) "Creation of an Adaptive Classifier to enhance the classification accuracy of existing classification algorithms in the field of Medical Data Mining," 2nd International Conference on Computing for Sustainable Global Development (INDIACom), New Delhi, 2015, pp. 376–381

4. Okfalisa I, Gazalba, Mustakim, Reza NGI, "Comparative analysis of k-nearest neighbor and modified k-nearest neighbor algorithm for data classification," 2017 2nd International conferences on Information Technology, Information Systems and, Engineering E (2017) (ICITISEE), Yogyakarta, pp. 294–298, doi: 10.1109/ICITISEE.2017.8285514

5. Pristyanto Y, Pratama I, Nugraha AF (2018) "Data level approach for imbalanced class handling on educational data mining multiclass classification," International Conference on Information and Communications Technology (ICOIACT), Yogyakarta, 2018, pp. 310–314, doi: 10.1109/ICOIACT.2018.8350792