Code for Benchmarking machine learning for cardiovascular disease classification across 46 countries with geographic and temporal transportability assessment
Abstract
Background: Machine-learning models for cardiovascular disease (CVD) are often evaluated with random participant splits, which may overstate performance when countries differ in prevalence, risk-factor distributions, healthcare access, and survey implementation. We benchmarked statistical and machine-learning classifiers for prevalent self-reported CVD with explicit assessment of class imbalance, calibration, explainability, and geographic and temporal transportability. Methods: We analysed harmonised WHO STEPwise approach to noncommunicable disease risk factor surveillance (STEPS) surveys conducted from 2013 to 2023. The cohort included 187,791 adults, including 16,305 CVD cases, from 46 countries. Logistic regression, elastic-net logistic regression, random forest, and XGBoost were evaluated using five-fold country-grouped internal-external cross-validation with survey, country-year, and class weighting. We assessed ROC AUC, precision-recall AUC (PR AUC), Brier score, calibration, threshold-dependent performance, SHAP importance, and earlier-to-later transportability. Results: Weighted CVD prevalence was 7.78%. XGBoost achieved ROC AUC 0.671, PR AUC 0.162, and Brier score 0.0692; logistic regression achieved 0.669, 0.154, and 0.0694. Country-level performance varied substantially. Removing hypertension and diabetes reduced ROC AUC by 0.038-0.040. Temporal logistic-regression ROC AUC was 0.634 in Mongolia, 0.579 in Uganda, and 0.670 in Viet Nam. Conclusions: Complex models produced only small pooled gains over logistic regression, whereas geographic and temporal heterogeneity was substantial. Responsible public-health AI benchmarking should prioritise minority-class performance, calibration, robustness, explainability, and transportability rather than pooled ROC AUC alone.