Explainable XGBoost model and nomogram for risk factor identification and risk prediction in cerebral small vessel disease: a machine learning-based retrospective cohort study
An interpretable XGBoost-based ML model that facilitates early risk stratification and targeted interventions for CSVD is validated, readily transferable to resource-limited settings and on embedding the nomogram into electronic-health-record decision support.
Abstract
Background Cerebral small vessel disease (CSVD) is a common, clinically significant vascular disorder that frequently leads to cognitive impairment, dementia, and poor overall prognosis. Owing to its complex hemodynamic characteristics and multifactorial pathophysiology, early identification of individuals at high risk for CSVD remains a clinical challenge. This study aimed to develop and validate an interpretable machine learning (ML) model for predicting the occurrence of CSVD. Methods We retrospectively enrolled 1,640 adult patients treated at the Fifth Affiliated Hospital of Xinjiang Medical University between September 2019 and December 2024. Twenty-three candidate variables (demographics, vitals, biomarkers, comorbidities) were evaluated. Feature selection was performed using least absolute shrinkage and selection operator (LASSO) regression, followed by stepwise backward elimination in multivariable logistic regression. Six supervised ML algorithms (DT, KNN, LR, LightGBM, XGBoost, SVM) were compared. Performance was assessed using ROC curves, calibration plots, and decision curve analysis (DCA). The optimal model was interpreted using SHapley Additive exPlanations (SHAP), and a bedside clinical nomogram was constructed. Results Ten independent predictors were identified: blood glucose, history of hypertension, systolic blood pressure, age, triglycerides, history of stroke, cystatin C, C-reactive protein, homocysteine, and body mass index. Among all models, XGBoost demonstrated the best performance, with an AUC of 0.968 in the training cohort and 0.938 in the validation cohort. Calibration plots and DCA confirmed its clinical utility. The derived nomogram demonstrated strong prognostic discrimination (p < 0.0001). The XGBoost model achieved an accuracy of 88.0%, sensitivity of 80.9%, specificity of 93.8%, and an F1 score of 0.86, corresponding to a 5.4-percentage-point gain in AUC over logistic regression. Ten-fold cross-validation confirmed this ranking, with a mean AUC of 0.934 ± 0.016. Conclusions We validated an interpretable XGBoost-based ML model that facilitates early risk stratification and targeted interventions for CSVD. Because the model relies only on routinely collected, low-cost variables and open-source software, it is readily transferable to resource-limited settings; future work will focus on prospective, multicentre external validation and on embedding the nomogram into electronic-health-record decision support.
Coronary artery aneurysm (CAA) is the most serious complication of Kawasaki disease (KD) and is associated with long-term cardiovascular morbidity. Early identification of children at high risk of CAA remains challenging. We aimed to develop and temporally validate a machine learning–based prediction model for individualized CAA risk assessment in children with KD.
In this retrospective cohort study, 1,930 children diagnosed with KD at The Second Affiliated Hospital of Wenzhou Medical University between January 2018 and March 2025 were included. Among them, 1,378 patients diagnosed between January 2018 and December 2023 constituted the development cohort and were randomly divided into a training set and an internal test set. A further 552 patients diagnosed between January 2024 and March 2025 were used as an independent temporal validation cohort. Forty-two demographic, clinical, laboratory, and composite biomarker variables were initially considered. Least Absolute Shrinkage and Selection Operator (LASSO) regression was used for feature selection after preprocessing. Ten machine-learning algorithms were developed and compared. The model was designed to be applied after the initial IVIG response had been determined. Model performance was evaluated using discrimination, calibration, and decision curve analysis. Model interpretability was assessed using Shapley Additive Explanations (SHAP).
12 predictors were retained in the final model, including age, intravenous immunoglobulin (IVIG) resistance, IVIG administration time, oral mucosal changes, albumin, hematocrit, platelet count, prothrombin time, D-dimer, NT-proBNP, prognostic nutritional index, and C-reactive protein–to–albumin ratio. Among the ten algorithms, the Random Forest model showed the best overall performance, achieving an area under the receiver operating characteristic curve (AUC) of 0.899 (95% CI, 0.870–0.928) in the internal test set and 0.847 (95% CI, 0.812–0.882) in the temporal validation cohort. Calibration and decision curve analyses indicated good agreement and clinical utility.
We developed and temporally validated a machine-learning model for predicting CAA risk in children with KD. The model has been deployed as an online tool to support early risk stratification. Further prospective multicenter validation is needed before routine clinical implementation.
An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN.
Juan Lv, Xi-Rui Wang, Zhengyi Zhang· Frontiers in Medicine· 0 citations
Background Patients with peripheral artery disease (PAD) face high postoperative mortality risks, necessitating precise risk stratification. While machine learning offers superior performance, its black-box nature limits clinical utility, and the prognostic value of the neutrophil-to-lymphocyte ratio (NLR) remains controversial. Methods A total of 610 surgically managed PAD patients were enrolled (median follow-up: 4 years) and randomly split into training (70%) and test (30%) sets. Six machine learning algorithms were constructed and optimized. Model performance was evaluated using area under the receiver operating characteristic curve (AUC) and decision curve analysis (DCA). The sHapley additive exPlanations (SHAP) were employed for model interpretation and visualizing nonlinear relationships. Results The random forest model achieved optimal performance (test set AUC = 0.814) with significant clinical net benefit. SHAP analysis identified age, prothrombin activity, and Rutherford classification as top predictors. Notably, while multivariate Cox regression failed to identify NLR as a linear predictor, SHAP dependence plots revealed a distinct nonlinear pattern: risk contribution increased sharply at low standardized NLR values before plateauing. Conclusion We established an interpretable random forest model for predicting postoperative mortality in PAD. By integrating SHAP analysis, this study validates the nonlinear prognostic significance of NLR and demonstrates how explainable ML can complement traditional statistics for individualized risk assessment.
Yi-Fei Li, Qiang Zhang, Wenxin Zhao et al.· Frontiers in Cardiovascular...· 0 citations
Patients with diabetic foot (DF) have a high risk of cardiovascular (CV) death, yet dedicated risk-prediction tools for this population are lacking. We developed and temporally validated an interpretable machine learning (ML) model for predicting CV death in patients with DF. This single-center retrospective cohort study included 2,835 patients admitted between February 2017 and May 2025. The development cohort comprised 2,325 patients, including 748 CV deaths, and was divided into training, internal validation, and held-out test sets; an independent temporal validation cohort included 510 patients, including 220 CV deaths. Nine supervised ML algorithms were compared using the area under the receiver operating characteristic curve (AUC). Extreme gradient boosting (XGBoost) showed the best overall performance. The optimal model, incorporating demographic and diabetes-related characteristics, routine laboratory parameters, and DF-specific features, achieved AUCs of 0.829 (95% confidence interval [CI]: 0.752-0.905) in internal validation, 0.844 (95% CI: 0.806-0.881) in the held-out test set, and 0.828 (95% CI: 0.789-0.868) in temporal validation. The model demonstrated good calibration and favorable net benefit on decision curve analysis. SHapley Additive exPlanations (SHAP) identified age, serum creatinine, glycated hemoglobin, triglycerides, and body mass index as the most influential predictors of increased model-predicted risk. This interpretable XGBoost model may support early identification and individualized risk stratification of patients with DF at high risk of CV death; however, prospective multicenter validation is required before clinical implementation.
Xiaoling Wan, Ting Shi, Qiao Liu et al.· Biomolecules & biomedicine· 0 citations
Background Deep vein thrombosis (DVT) is a common thrombotic condition with substantial morbidity when not identified early. Machine learning (ML)–based predictive models may improve early identification of patients at high risk for DVT, but few clinically applicable early-risk models exist. Objectives To develop and internally validate a ML model using routinely available clinical and laboratory indicators for early risk prediction of DVT, and to identify the most influential predictors using model explainability techniques. Methods We retrospectively analyzed clinical data from 231 patients evaluated at the Fifth Affiliated Hospital of Southern Medical University between January 2017 and June 2024. Patients were labeled as DVT occurrence (n = 159) or non-occurrence (n = 72). Seven candidate predictors were selected by Least Absolute Shrinkage and Selection Operator (LASSO) regression. The dataset was split into training (70%, n = 162) and test (30%, n = 69) sets. Five ML algorithms were trained: XGBoost, CatBoost, Random Forest (RF), Logistic Regression, and Support Vector Machine, with hyperparameter tuning on the training set. Model performance was assessed by 5-fold cross-validation and on the held-out test set using Area Under the Receiver Operating Characteristic Curve (AUC), accuracy, recall, and F1 score. The best model was further interpreted via feature importance and Shapley Additive Explanations (SHAP). Results LASSO selected seven predictors: hemoglobin, platelet count, leukocyte count, fibrinogen, prothrombin time, D-dimer (DD), and glucose. The Random Forest model showed the best discrimination (test-set AUC = 0.874), with favorable accuracy, recall, and F1 compared with other classifiers (detailed metrics reported in the manuscript). In the RF model, D-dimer had the highest feature-importance contribution; SHAP analysis confirmed DD as the dominant risk driver and characterized the directions and relative effects of other features. Conclusions We developed an internally validated ML model for early DVT risk prediction using seven routine clinical variables; Random Forest achieved the best performance and identified D-dimer as the most influential predictor. This model may support earlier identification and intervention for patients at risk of DVT, pending external validation and prospective evaluation.
Xue Wang, Xiankai Chen, Jun Mao et al.· PeerJ· 0 citations
BACKGROUND
Cardiovascular disease (CVD) is a major concern among cancer survivors. However, the intersection of cancer and CVD has only recently gained broader attention, and substantial evidence gaps remain. This study aimed to identify risk factors associated with incident CVD in cancer survivors and to develop a machine learning model for CVD risk prediction.
METHODS
In this retrospective study, we included 2,500 patients receiving systemic antitumor therapy at Jilin Cancer Hospital; 188 incident CVD events were observed. Variables spanning demographic, clinical, tumor- and treatment-related, and laboratory domains were collected. The dataset was randomly split into a training set (75%) and an internal validation set (25%). Feature selection was performed using LASSO regression. Six prediction models were developed, including five machine learning algorithms (GBM, CoxBoost, XGBoost, SVM, and RSF) and a traditional Cox proportional hazards model. Model performance was evaluated using time-dependent receiver operating characteristic curves (AUC), calibration analyses, and decision curve analysis. Model interpretability was assessed using SHapley Additive exPlanations (SHAP).
RESULTS
In univariate analyses, 48 variables were associated with CVD risk (P < 0.20). LASSO regression identified 19 predictors for model development. Key predictors included elevated systolic blood pressure, specific cancer types, anthracycline use, and a history of hypertension. The XGBoost model demonstrated the best predictive performance, with an average AUC of 0.666, sensitivity of 69.8%, specificity of 60.7%, and overall accuracy of 61.4% in the internal validation set. The model showed good calibration and yielded a positive net benefit across a range of clinical thresholds. SHAP analysis indicated that cancer type, lack of anti-HER2 therapy, elevated systolic blood pressure, advanced T stage, advanced TNM stage, and higher uric acid levels were the most influential predictors.
CONCLUSION
This study developed and internally validated interpretable machine learning models to predict CVD risk among cancer survivors. The models demonstrated good discrimination and calibration and outperformed traditional methods. By enabling individualized risk quantification and providing transparent interpretation of key predictors, this approach offers a practical tool to support personalized surveillance and prevention strategies in cardio-oncology.