Comparing Different Oversampling Methods in Predicting Multi-Class Educational Datasets Using Machine Learning Techniques

Tariq Muhammad Arham; Sargano Allah Bux; Iftikhar Muhammad Aksam; Habib Zulfiqar

doi:10.2478/cait-2023-0044

Cybernetics and Information Technologies (Nov 2023)

Comparing Different Oversampling Methods in Predicting Multi-Class Educational Datasets Using Machine Learning Techniques

Tariq Muhammad Arham,
Sargano Allah Bux,
Iftikhar Muhammad Aksam,
Habib Zulfiqar

Affiliations

Tariq Muhammad Arham: 1University of Central Punjab, Department of Computer Science, Lahore, Pakistan
Sargano Allah Bux: 2COMSATS University Islamabad, Department of Computer Science, Lahore, Pakistan
Iftikhar Muhammad Aksam: 2COMSATS University Islamabad, Department of Computer Science, Lahore, Pakistan
Habib Zulfiqar: 2COMSATS University Islamabad, Department of Computer Science, Lahore, Pakistan

DOI: https://doi.org/10.2478/cait-2023-0044
Journal volume & issue: Vol. 23, no. 4
pp. 199 – 212

Abstract

Read online

Predicting students’ academic performance is a critical research area, yet imbalanced educational datasets, characterized by unequal academic-level representation, present challenges for classifiers. While prior research has addressed the imbalance in binary-class datasets, this study focuses on multi-class datasets. A comparison of ten resampling methods (SMOTE, Adasyn, Distance SMOTE, BorderLineSMOTE, KmeansSMOTE, SVMSMOTE, LN SMOTE, MWSMOTE, Safe Level SMOTE, and SMOTETomek) is conducted alongside nine classification models: K-Nearest Neighbors (KNN), Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Support Vector Machine (SVM), Logistic Regression (LR), Extra Tree (ET), Random Forest (RT), Extreme Gradient Boosting (XGB), and Ada Boost (AdaB). Following a rigorous evaluation, including hyperparameter tuning and 10 fold cross-validations, KNN with SmoteTomek attains the highest accuracy of 83.7%, as demonstrated through an ablation study. These results emphasize SMOTETomek’s effectiveness in mitigating class imbalance in educational datasets and highlight KNN’s potential as an educational data mining classifier.

Published in Cybernetics and Information Technologies

ISSN: 1314-4081 (Online)
Publisher: Sciendo
Country of publisher: Poland
LCC subjects: Science: Science (General): Cybernetics
Website: https://sciendo.com/journal/cait

About the journal

Abstract

Keywords