Skip to content
Review Open access

An explainable machine learning framework for early student dropout risk prediction and stratification

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 45 references

Abstract

Student dropout remains a persistent challenge in higher education institutions, affecting academic continuity, institutional performance, and long-term socioeconomic outcomes. Early identification of at-risk students enables timely intervention and improved retention strategies. This study proposes an explainable machine learning framework for identifying and stratifying students according to a constructed dropout-risk label using academic, behavioural, financial, health-related, and personal factors. A dataset of 253 student records was collected through structured Google Form surveys and transformed using ordinal encoding into machine-learning-ready features, including CGPA, attendance, stress level, financial difficulty, health issues, course interest, thoughts of discontinuation, and personal or family problems. Three classification models Logistic Regression, Decision Tree, and Random Forest were trained and evaluated using a stratified 70:30 train-test split. Performance was assessed using Accuracy, Precision, Recall, F1-score, and ROC-AUC. Experimental results demonstrated that the Random Forest model achieved the best overall performance with an accuracy of 0.86, precision of 0.96, recall of 0.88, F1-score of 0.92, and an AUC score of 0.911. The system integrates probability-based risk categorization and an explainability module within a Flask-based web prototype dashboard, providing a transparent and practical decision-support tool for educational institutions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.