Skip to content
Open access

AI-Powered Student Dropout Prediction and Personalized Intervention Using TC-Net in Education

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 40 references

TL;DR

The findings demonstrate the capability of AI models in minimising dropout risks and enhancing academic performance through prompt, data-informed assistance.

Abstract

Increasing student dropout rates, caused by social, academic, and personal obstacles, raise significant concerns for educational systems. This research introduces an AI-based method for the early detection and support of at-risk students utilizing the Student Final Grade Prediction dataset from two schools in Portugal. Following data preprocessing, which involved cleaning and feature extraction via t-distributed Stochastic Neighbor Embedding (t-SNE), a hybrid model integrating Tabular Data Network and Capsule Networks (TC-Net) was created for performance forecasting. The model obtained a Mean Squared Error (MSE) of 0.43 and an R-squared value of 60.57%, demonstrating robust predictive accuracy. Personalised strategies were then implemented, targeting 90% of at-risk students. Specifically, 80% participated in tailored learning plans, and 70% accessed tutoring support. The findings demonstrate the capability of AI models in minimising dropout risks and enhancing academic performance through prompt, data-informed assistance.

Read PDF

Similar papers

Open access Sep 2026

Interpretable machine learning approaches for student dropout prediction in higher education

Student dropout remains one of the most significant challenges in higher education, affecting academic performance, financial sustainability, and strategic planning within universities. This study presents an approach to predicting student dropout risk using machine learning methods and educational analytics. The research is based on an open-access dataset from the UCI Machine Learning Repository containing 4,424 records and 36 attributes without missing values. Three experimental datasets were constructed using different preprocessing strategies, including class balancing, logarithmic feature transformation, and processing of the Enrolled category. Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), k-Nearest Neighbours (k-NN), and Neural Network (NN) models were implemented and compared. The results showed that Logistic Regression and Random Forest achieved the highest performance, with accuracy above 90% and ROC AUC values of up to 0.96. The most significant risk factors included academic performance during the first and second semesters, tuition fee payment status, outstanding debt, scholarship status, age at enrolment, and study programme. The findings indicate that machine learning methods can effectively support early dropout prediction and decision-support systems in higher education institutions.

Arūnas Mincevičius · 0 citations
Open access 2026

Hybrid Approach Based on Supervised Learning and Synthetic Data Generation for University Student Dropout Risk Prediction

: Student dropout in higher education constitutes a structural problem that affects academic quality and institutional sustainability. In Colombia, between 30% and 50% of students left their studies, highlighting the need to strengthen early detection systems. However, the performance of supervised models is often affected by class imbalance, data scarcity, and constraints associated with the use of sensitive information. This study proposes a hybrid methodological approach structured under the CRISP-DM framework that integrates a supervised learning model with synthetic data generated by generative artificial intelligence. The workflow incorporates exploratory data analysis, business rules embedded in the generative process, statistical validation of synthetic data, and comparative evaluation under a no-data-leakage setting. The results show that incorporating synthetic data exclusively into the training set improves predictive performance, particularly for minority classes, reflected in gains in precision, recall, and F1-score, while preserving evaluation on untouched real test data. The proposed approach offers a reproducible, transferable methodological alternative to mitigate class imbalance in university student dropout prediction.

L. Caicedo, Juan Muñoz, N. Díaz · 0 citations
Conference Jul 2026

Early Student Dropout Prediction Using Machine Learning

Student dropout is a significant issue in higher education, affecting both students and institutions. Early identification of at-risk students can help universities improve retention. This study addresses student dropout prediction as a binary classification problem using 4,424 student records. To support realistic early prediction, only features available during the early stages of academic study were used. Three machine learning models were evaluated: Logistic Regression, Random Forest, and XGBoost. Class imbalance was handled through class weighting and parameter adjustment, while stratified 10-fold cross-validation ensured result stability. Experimental results showed that Random Forest achieved the highest accuracy (0.866), whereas XGBoost obtained the best recall (0.81) and ROC-AUC (0.913), making it particularly effective for identifying at-risk students. Logistic Regression provided a reliable and interpretable baseline. Feature importance analysis revealed that academic performance and financial factors were among the strongest predictors of dropout. The findings demonstrate the potential of machine learning as an early warning system for educational decision-making. However, the study is limited to a single dataset and does not include temporal data. Future work may incorporate additional features and evaluate model performance across different institutions.

Saeed Al Sagherji, Rania Alhalaseh, Mohammad Abbadi · 0 citations
Review Open access Jul 2026

Academic Risk Profiling and Tailored Student Guidance Through Machine Learning in Kinshasa Universities

Educational data can help identify, before the end of a semester, academic trajectories that deserve timely attention. This study develops an academic risk profiling and tailored student guidance using machine learning techniques for universities in Kinshasa. A quantitative, experimental, and predictive design was applied to a harmonized dataset of 9,000 student-semester observations and fourteen predictors after removing a redundant composite engagement index. Data preparation, modelling, and visualisation were performed in a reproducible Python/Jupyter Notebook workflow using pandas, NumPy, scikit-learn, and Matplotlib. A stratified 80/20 training-test split and five-fold stratified cross-validation were used to compare multinomial logistic regression, decision tree, random forest, and multilayer perceptron models. On the independent test set, the multilayer perceptron achieved the highest accuracy (0.720) and macro-AUC-ROC (0.945), while logistic regression achieved the highest balanced accuracy (0.714). Random forest was retained because it achieved the highest macro F1 score (0.686), the prespecified criterion for protecting attention to minority classes, together with macro precision of 0.713 and macro One-vs-Rest AUC-ROC of 0.941. Continuous assessment average, midterm average, and the previous validated-credit rate were the most informative signals. Recommendations are derived from the predicted class, while uncertain cases are flagged for human review. The proposed model is therefore a decision-support tool rather than an automated decision-maker.

Augustin Pambi Tadiamba, Pierre K. Kafunda, David M. Kutangila et al. · 0 citations
Open access Jul 2026

A Proposed Hybrid Deep Learning and Ensemble Learning Model for Predicting Student Dropout in Education

Student dropout remains a major challenge for higher education institutions due to its academic, social, and economic consequences. Early identification of students at risk of dropping out is essential for implementing timely interventions; however, existing prediction approaches often rely on standalone machine learning models that inadequately capture the complex, nonlinear relationships within heterogeneous educational data and exhibit reduced performance under class imbalance. To address these limitations, this study proposes a hybrid prediction framework that integrates deep neural representation learning with a stacked ensemble architecture to improve the accuracy and robustness of student dropout prediction. The proposed framework is evaluated on three benchmark educational datasets. A deep neural network is employed to learn latent feature representations from preprocessed student data, and the resulting embeddings are used to train heterogeneous machine learning classifiers, including XGBoost, Decision Tree, AdaBoost, Extra Trees, and K-Nearest Neighbors. Their probabilistic predictions are subsequently combined through a logistic regression meta-learner to generate the final prediction. To investigate the impact of class imbalance, both class weighting and the Synthetic Minority Over-sampling Technique (SMOTE) are incorporated and systematically evaluated. Model performance is assessed using accuracy, balanced accuracy, precision, recall, F1-score, and ROC-AUC, while permutation feature importance is employed to examine the influence of the original input variables on the overall prediction pipeline, and McNemar's test is used to assess the statistical significance of performance differences. Experimental results demonstrate that the proposed framework consistently outperforms standalone machine learning models, a standalone deep neural network, and conventional stacking approaches across all three datasets. These findings demonstrate that combining deep representation learning with stacked ensemble learning provides a robust and effective framework for early student dropout prediction, supporting data-driven educational interventions and informed decision-making in higher education.

Abdulalim M. Ibrahim, M. Abonazel, Abdul-Hadi N. Ahmed · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.