Effect of Data Imbalance in Predicting Student Performance in a Structural Analysis Graduate Attribute-Based Module Using Random Forest Machine Learning
Masikini Lugoma, Abel Omphemetse Zimbili, Masengo Ilunga, Ngaka Mosia, Agarwal Abhishek
Authors Information |
Citation |
Full Text |
Masikini Lugoma
Department of Mining, Minerals and Geomatics Engineering, University of South Africa, Pretoria, South Africa
Abel Omphemetse Zimbili
Department of Civil and Environmental Engineering and Building Science, University of South Africa, Pretoria, South Africa
Masengo Ilunga
Department of Civil and Environmental Engineering and Building Science, University of South Africa, Pretoria, South Africa
Ngaka Mosia
Department of Industrial Engineering and Engineering Management, University of South Africa, Pretoria, South Africa
Agarwal Abhishek
Department of Mechanical Engineering, College of Science and Technology, Royal University of Bhutan, Phuentsholing, Bhutan
Cite this paper as:Lugoma, M., Zimbili, A. O., Ilunga, M., Mosia, N., Abhishek, A. (2025). Effect of Data Imbalance in Predicting Student Performance in a Structural Analysis Graduate Attribute-Based Module Using Random Forest Machine Learning.
Journal of Systemics, Cybernetics and Informatics, 23(2), 15-22. https://doi.org/10.54808/JSCI.23.02.15
Online ISSN (Journal): 1690-4524
Abstract
This study uses Random Forest algorithm to model students' final year mark in an engineering technology module taught by the University of South Africa. The algorithm uses a supervised learning classification technique to map the different assessment marks and the final mark. Hence, the latter are labelled instances whereas the former constitute the features. Random Forest (RF) has been applied to Structural Analysis 3, which takes into consideration the graduate attribute concept or level of competence as far as assessments are concerned. Firstly, the RF is subjected to imbalanced binary classes, then balanced classes are achieved by Synthetic Minority Oversampling Technique (SMOTE) and class weights adjustment techniques. The results showed that SMOTE brought an improvement in accuracy of 3%. It was also revealed that an increase of 4, 15 and 9% in precision, recall and F1-Score were observed in predicting non-competent students. An increase of 4 and 3% was noticed in the case of the precision and F1-Score respectively in predicting competent students, whereas the recall did not display any change. Despite the RF with SMOTE overperformed standard RF and RF class weights adjustment, all three algorithms were good candidates in the prediction of student performance. RF-SMOTE could be suggested as a guiding instrument when dealing with imbalanced data.