Skip to content
Open access

Explainable Machine Learning-Based Decision Support for Diabetes Risk Screening Using Public Health Indicator Data

Aug 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 27 references

Abstract

Diabetes remains a major public-health challenge because many individuals at elevated risk are identified only after avoidable metabolic deterioration or complication pathways have already begun. This study develops and evaluates an explainable machine-learning decision-support framework for diabetes risk screening using the Kaggle Diabetes Health Indicators Dataset derived from the 2015 Behavioral Risk Factor Surveillance System (BRFSS). The framework is designed for screening support and population-level risk stratification rather than clinical diagnosis. It combines data quality assessment, class-imbalance-aware modelling, clinically meaningful feature engineering, stratified splitting, supervised learning, threshold analysis, calibration assessment, subgroup reliability analysis, and explainability using SHAP and LIME. The main modelling protocol focuses on Logistic Regression, Decision Tree, Random Forest, and XGBoost, while additional benchmark outputs from AdaBoost, Naive Bayes, and K-Nearest Neighbors are retained as exploratory comparators from the supplied notebook. The notebook audit showed an apparent high KNN hold-out performance; however, cross-validation and balanced-file sensitivity analysis indicated that Random Forest and XGBoost provided more stable and methodologically defensible behaviour for screening-oriented deployment. SHAP analysis identified general health, comorbidity burden, age, BMI, high blood pressure, sex, income, and high cholesterol as leading model contributors. Threshold analysis showed that lowering the decision threshold increased sensitivity and reduced false negatives, which is important in screening applications. The contribution of this work is an interpretable and reproducible public-health decision-support pipeline that links predictive performance with explanation, calibration, threshold selection, and subgroup reliability.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.