Implementasi Prediksi Stunting Menggunakan Algoritma Xgboost dan Smote Berdasarkan Data BKKBN
Abstract
The management and intervention of stunting cases in toddlers still face major challenges due to the high class imbalance in risk assessment data, where the population of normal toddlers significantly dominates the at-risk class. This condition often causes machine learning classification models to suffer from prediction bias and reduced accuracy. This study aims to build a Machine Learning classification model based on the XGBoost algorithm combined with the Synthetic Minority Over-sampling Technique (SMOTE) to predict stunting risk and analyze its primary key determinants (Feature Importance) using Stunting Risk Family (KRS) data from BKKBN Central Java. The dataset comprises 7,605 family data records with evaluation performed on 2,771 pure test data points. The evaluation results demonstrate that the combination of XGBoost and SMOTE achieved an Accuracy of 88.20% and a Precision of 99.91%. Furthermore, Gain analysis confirms that Gender (score 0.3232) and Sanitation/Latrine Facilities (score 0.2182) are the two most dominant risk factors for stunting. These findings demonstrate the potential of leveraging artificial intelligence systems to support efficient decision-making for healthcare workers within a smart governance framework.