An ensemble model that combines three ML algorithms Random Forest, K-nearest Neighbors, and ADABOOST is proposed that is higher than the accuracies of the compared models and integrated through a voting mechanism.
Student academic performance is an important factor in evaluating the effectiveness of the learning process and
identifying students who may require additional academic support. Traditional methods of evaluating student performance
mainly depend on examination marks and teacher observations, which may not provide sufficient information for early
identification of students at academic risk. This paper presents a machine learning-based approach for predicting student
performance using relevant academic and personal attributes. The proposed system involves data preprocessing, feature
selection, model training, and performance evaluation. Machine learning algorithms such as Linear Regression, Decision Tree,
Random Forest, and Support Vector Machine can be applied to identify patterns in student data and predict their expected
academic performance. The performance of the models can be compared using suitable evaluation metrics such as accuracy,
Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE), depending on the prediction
task. The proposed approach can assist educational institutions in identifying students who may need additional guidance and
support. The system demonstrates the potential of machine learning in supporting data-driven academic decision-making.
S. S, S. R, P. R. et al.· International Journal for Re...· 0 citations
The present paper explores some of the most important determinants of academic achievement and assesses several predictive models based on the Student Performance Factors (SPF) dataset and implies that the monitoring of attendance should become the central element of any academic early warning system.
Shang-Jia Wang· Mathematical Modeling and Al...· 0 citations
Student performance prediction has become an important research area in educational data mining because it
enables educational institutions to identify academically at-risk students and implement timely intervention strategies.
Although machine learning techniques have significantly improved prediction accuracy, many existing models function as
black-box systems that provide limited explanation of the factors influencing academic performance. The absence of
interpretability restricts educators from understanding the reasoning behind prediction outcomes and reduces confidence
in adopting artificial intelligence-based decision support systems. This study proposes an explainable ensemble learning
approach for student performance prediction by integrating Random Forest, XGBoost, and LightGBM classifiers with
SHapley Additive exPlanations (SHAP). The proposed framework includes data preprocessing, feature engineering,
ensemble learning, and feature interpretation to improve both predictive performance and model transparency. A publicly
available student performance dataset containing 14,003 student records with 16 academic, behavioural, and demographic
attributes is used for experimental evaluation. The dataset is divided into training and testing subsets using an 80:20 ratio.
The performance of the proposed approach is evaluated using Accuracy, Precision, Recall, F1-Score, and ROC-AUC and
compared with conventional machine learning models. SHAP analysis is employed to identify the contribution of
individual features influencing student performance, enabling transparent and interpretable predictions. The proposed
approach assists educators in identifying the key factors affecting academic achievement and supports timely intervention
for improving student success. The results demonstrate that combining ensemble learning with explainable artificial
intelligence provides an effective and reliable framework for educational decision-making.
S. P· International Journal of Inn...· 0 citations
Predicting secondary school students' academic accomplishments is crucial for early intervention and personalized learning strategies. This study develops a machine learning-based system to forecast student performance, including grade and percentage prediction, while analyzing the impact of various socio-economic, educational, personal, and technological factors. The dataset was collected through a structured survey, incorporating aspects such as family income, parental education, access to private tutoring, school infrastructure, learning environment, mental health, career guidance, geographic constraints, government policies, and digital literacy. Pre-processing was done using five machine learning algorithms: Random Forest, Support Vector Machine, Decision Tree, K-Nearest Neighbors, and Gradient BoostingTo assess model performance, various evaluation metrics, such as accuracy and root mean squared error, were utilized. The findings suggest that machine learning methods are capable of accurately forecasting student performance, with the Random Forest algorithm demonstrating the greatest level of precision. This study lays the groundwork for AI-based educational resources aimed at recognizing students who are at risk and facilitating focused interventions.
Shravani P.Pawar, S. Deshmukh, Priya Chandran· Enterprise Development and M...· 0 citations
: Amid the global acceleration of digital transformation in education, achieving precision teaching and improving student learning outcomes have become central concerns in the educational sector. This study explores the application of machine learning models — Random Forest and Support Vector Machine — in predicting student academic performance. Using an open-source dataset from Kaggle, the research selects 13 behavioral and demographic indicators, including study hours, mental health, attendance, and lifestyle habits. The results show that both random forest and support vector machine achieved high accuracy, but RF demonstrated better recall for identifying at-risk students, making it more suitable for early intervention systems. Feature importance analysis reveals that daily study time and mental health ratings are the most influential predictors. RF integrates a wider range of behavioral traits, while SVM relies more heavily on entertainment-related variables, leading to lower robustness. Visualization of prediction results enhances interpretability and supports data-driven educational decisions. The study concludes by emphasizing the need to incorporate broader psychological and social factors in future models to improve prediction generalization and provide actionable insights for personalized teaching strategies.
Chenhao Sun, Hewen Sun· Proceedings of the 3rd Inter...· 0 citations
This paper compares three machine learning algorithms—k-Nearest Neighbours (k-NN), Random Forest, as well as Support Vector Machine (SVM)—for predicting high school student performance, using actual exam data from 12,211 students in Jorhat, Assam. The most important decision was to keep all student records (109 absent students and 2 withheld results) instead of deleting them. This methodology made the models more realistic and applicable to real schools. The results demonstrate that Random Forest achieved the highest accuracy at 69.63%, followed by k-NN at 58.08%, and SVM at 57.35%. The total percentage of the grand total and minimum subject mark emerged as the most accurate predictors of student success. Current research demonstrates that it is now possible to identify struggling students with high accuracy, enabling timely interventions that support academic success.
Rinku Mani Kalita, Dr. Siddhartha Baruah· Journal of Intelligent Decis...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.