Skip to content
Open access

Early Prediction of Student Dropout Through Machine Learning in Educational Data Mining: A Comparative Analysis of Algorithms and Data Balancing Techniques

Sep 2026 · Applied Sciences · 0 citations · 25 references

Abstract

This research presents the development of a predictive model to address student dropout in higher education, using advanced artificial intelligence techniques, specifically in the field of machine learning. To this end, the CRISP-DM methodology is followed, and a set of administrative and academic data is analyzed, applying feature engineering and the SMOTE technique to balance the classes. The comparative analysis reveals that the optimized Random Forest algorithm is the most robust, achieving an accuracy of 84.09% and an AUC of 0.92. In addition, an intelligent classification analysis is performed to identify trends among students who may eventually drop out. The results obtained from this study are compared with those of a study conducted in Iceland. Through this comparison, it can be observed that academic performance and financial status are the main common indicators. An important contribution is the identification of so-called “risk profiles” through unsupervised learning using the K-Means algorithm, enabling early and personalized interventions that can prevent students from dropping out of their degree programs. This approach not only improves retention at institutions but also establishes a technical standard that can be used to strengthen institutional educational policies based on the findings obtained.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.