Aug 2026· Mathematical Modeling and Algorithm Application· Vol 9, pp. 147-153· 0 citations· 11 references
TL;DR
The present paper explores some of the most important determinants of academic achievement and assesses several predictive models based on the Student Performance Factors (SPF) dataset and implies that the monitoring of attendance should become the central element of any academic early warning system.
Abstract
The fast growth of educational data systems has led to more student data becoming available at scale to use in learning analytics. It is important to effectively analyze these data to forecast academic performance, as this will facilitate early detection of risks and individualized interventions. The present paper explores some of the most important determinants of academic achievement and assesses several predictive models based on the Student Performance Factors (SPF) dataset (N = 6,607). Linear Regression (LR) and Random Forest (RF) models are built and benchmarked against each other. As can be seen, the RF model, according to the results (R² = 0.70), is significantly outperforming the LR model (R² = 0.62), and the error rate has been minimized by 11.5 percent. Feature importance analysis indicates that attendance is the main determinant (importance weight = 0.381), followed by study hours (0.243) and past scores (0.091). Interestingly, the combination of the existing learning behaviors is six times greater than the historical performance, and the family background factors have insignificant direct effects. The results obtained can be used to justify data-driven educational interventions and imply that the monitoring of attendance should become the central element of any academic early warning system.
Student academic performance is an important factor in evaluating the effectiveness of the learning process and
identifying students who may require additional academic support. Traditional methods of evaluating student performance
mainly depend on examination marks and teacher observations, which may not provide sufficient information for early
identification of students at academic risk. This paper presents a machine learning-based approach for predicting student
performance using relevant academic and personal attributes. The proposed system involves data preprocessing, feature
selection, model training, and performance evaluation. Machine learning algorithms such as Linear Regression, Decision Tree,
Random Forest, and Support Vector Machine can be applied to identify patterns in student data and predict their expected
academic performance. The performance of the models can be compared using suitable evaluation metrics such as accuracy,
Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE), depending on the prediction
task. The proposed approach can assist educational institutions in identifying students who may need additional guidance and
support. The system demonstrates the potential of machine learning in supporting data-driven academic decision-making.
S. S, S. R, P. R. et al.· International Journal for Re...· 0 citations
Student performance prediction has become an important research area in educational data mining. This project proposes a Machine Learning–based system that analyzes academic and behavioral data to predict student performance in advance. The system uses algorithms such as Decision Tree, Random Forest, and Logistic Regression to classify students into performance categories (High, Medium, and Low) or predict final grades. The dataset includes attributes such as attendance, internal marks, assignment scores, study hours, and previous semester results. The trained model identifies at-risk students early and provides actionable insights for teachers and institutions. The proposed system improves academic monitoring, enables early intervention, and enhances overall student success rates through data-driven decision-making.
Shaheela Y, Vigashini S, Surya Prakash Av et al.· International Journal Of Rec...· 0 citations
: Amid the global acceleration of digital transformation in education, achieving precision teaching and improving student learning outcomes have become central concerns in the educational sector. This study explores the application of machine learning models — Random Forest and Support Vector Machine — in predicting student academic performance. Using an open-source dataset from Kaggle, the research selects 13 behavioral and demographic indicators, including study hours, mental health, attendance, and lifestyle habits. The results show that both random forest and support vector machine achieved high accuracy, but RF demonstrated better recall for identifying at-risk students, making it more suitable for early intervention systems. Feature importance analysis reveals that daily study time and mental health ratings are the most influential predictors. RF integrates a wider range of behavioral traits, while SVM relies more heavily on entertainment-related variables, leading to lower robustness. Visualization of prediction results enhances interpretability and supports data-driven educational decisions. The study concludes by emphasizing the need to incorporate broader psychological and social factors in future models to improve prediction generalization and provide actionable insights for personalized teaching strategies.
Chenhao Sun, Hewen Sun· Proceedings of the 3rd Inter...· 0 citations
The current research provides a thorough exploration into different methods of machine learning used to predict educational performance using a variety of data sources. This research studies methods for predicting academic performance and displays the difference in performance of each model, including performance measures, demographics, and behavior survey data collected from middle school-aged children. In this research, 5-fold cross-validation was useful for demonstrating increases in the accuracy of predictions with the use of multiple types of data without significantly affecting the level of computing needed for making predictions. The research showed that using multiple data types significantly increases the predictive power of the model. Additionally, among all methods evaluated, multiple linear regression had the best balance between accuracy and computational requirement for predicting student performance. The factors determined to be the most predictive of student academic performance included student behavior, the educational level of parents, socio-economic status, and historically earned grades. The study of educational data mining presented in this paper provides a valuable understanding of how different data integration configurations affect how well algorithms compare in performance. The results may be used as criteria for schools seeking to employ data-driven strategies for improving student achievement and academic performance.
Botan Onat, A. Bilge, A. Akın· International journal of 3d...· 0 citations
This paper compares three machine learning algorithms—k-Nearest Neighbours (k-NN), Random Forest, as well as Support Vector Machine (SVM)—for predicting high school student performance, using actual exam data from 12,211 students in Jorhat, Assam. The most important decision was to keep all student records (109 absent students and 2 withheld results) instead of deleting them. This methodology made the models more realistic and applicable to real schools. The results demonstrate that Random Forest achieved the highest accuracy at 69.63%, followed by k-NN at 58.08%, and SVM at 57.35%. The total percentage of the grand total and minimum subject mark emerged as the most accurate predictors of student success. Current research demonstrates that it is now possible to identify struggling students with high accuracy, enabling timely interventions that support academic success.
Rinku Mani Kalita, Dr. Siddhartha Baruah· Journal of Intelligent Decis...· 0 citations
This research presents the development of a predictive model to address student dropout in higher education, using advanced artificial intelligence techniques, specifically in the field of machine learning. To this end, the CRISP-DM methodology is followed, and a set of administrative and academic data is analyzed, applying feature engineering and the SMOTE technique to balance the classes. The comparative analysis reveals that the optimized Random Forest algorithm is the most robust, achieving an accuracy of 84.09% and an AUC of 0.92. In addition, an intelligent classification analysis is performed to identify trends among students who may eventually drop out. The results obtained from this study are compared with those of a study conducted in Iceland. Through this comparison, it can be observed that academic performance and financial status are the main common indicators. An important contribution is the identification of so-called “risk profiles” through unsupervised learning using the K-Means algorithm, enabling early and personalized interventions that can prevent students from dropping out of their degree programs. This approach not only improves retention at institutions but also establishes a technical standard that can be used to strengthen institutional educational policies based on the findings obtained.
Edwin Guamán-Hidalgo, L. Enciso· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.