Back to feed
Open access

Predictive modeling of early diabetes diagnosis: An evaluation of XGBoost, support vector machine, and random forest classifiers

Jul 2026 · International Journal of Science and Research Archive · 0 citations

Abstract

This study addresses the challenge of delayed diagnosis of diabetes, a condition that often leads to severe complications if not detected early. The primary objective is to evaluate and compare the performance of three machine learning classifiers XGBoost, Support Vector Machine (SVM), and Random Forest for early diabetes prediction using clinical and lifestyle data. The study utilizes the Diabetes Health Indicators dataset, which includes features such as body mass index (BMI), blood pressure, cholesterol levels, and physical activity. The dataset was sourced from a publicly available repository and preprocessed through handling missing values, feature scaling, and encoding categorical variables. The models were trained on the processed dataset and evaluated using accuracy, precision, recall, and F1-score metrics, alongside exploratory data analysis to understand feature relationships. Results show that all three models performed effectively, with XGBoost achieving the highest accuracy of 85.11%, followed by SVM at 84.82%, and Random Forest at 83.16%. These findings highlight the strength of ensemble and boosting techniques in handling complex health data and accurately predicting diabetes risk. In conclusion, machine learning models demonstrate strong potential for supporting early diabetes diagnosis and improving clinical decision-making. It is recommended that healthcare systems adopt XGBoost-based predictive models in clinical decision support tools for early screening, while future studies should validate these models using real-world clinical data to enhance reliability and generalizability.

Read PDF