Comparative Evaluation of Machine Learning Models for Early Prediction of Student Academic Performance Using the Open University Learning Analytics Dataset
DOI:
https://doi.org/10.3991/ijim.v20i17.62574Keywords:
learning analytics; student performance prediction; machine learning; XGBoost; class imbalance; OULAD; educational data mining; early-warning systems.Abstract
Early identification of students at risk of academic failure is a central problem in learning analytics, yet the multi-table, behaviorally heterogeneous, and class-imbalanced nature of educational data complicates reliable prediction. This study develops and comparatively evaluates four supervised machine learning models—logistic regression, support vector machine (SVM) with a radial basis function kernel, random forest, and extreme gradient boosting (XGBoost)—for multi-class prediction of final student outcome on the Open University Learning Analytics Dataset (OULAD), comprising 32,593 anonymized student records and more than ten million virtual-learning-environment interaction logs. A reproducible, leakage-controlled feature-engineering pipeline is proposed: final examination records are excluded, and virtual-learning-environment activity is restricted to the first 120 days of each presentation, yielding 85 demographics, assessment, and behavioral features. All models are trained on an identical stratified 80/20 split and assessed within a unified evaluation framework using accuracy, weighted and macro F1-score, balanced accuracy, per-class recall, and stratified cross-validation. XGBoost achieved the strongest overall performance (accuracy 0.7302, macro F1 0.6658, balanced accuracy 0.6555, and cross-validated accuracy 0.7314), followed closely by random forest and logistic regression, while SVM, despite the weakest aggregate scores, produced the highest recall for the practically critical Fail class (0.4026). The analysis demonstrates that aggregate accuracy substantially overstates the usefulness of these models for risk detection and that the fail class remains the principal bottleneck. A deployed Streamlit prototype illustrates operationalization. The contribution lies in the leakage-aware early-window pipeline and the imbalance-sensitive, algorithm-aware comparison that together expose where predictive value is gained and lost.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Assel Omarbekova, Aizhan Nazyrova, Gulmira Bekmanova, Zhanar Lamasheva, Rakhila Turebayeva, Zhainagul Abzal

This work is licensed under a Creative Commons Attribution 4.0 International License.

