An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN.
Abstract
Background Patients with coexisting type 2 diabetes mellitus (T2DM) and hypertension (HTN) face a synergistically elevated risk of major adverse cardiovascular events (MACE). Evidence for prediction models developed specifically in established T2DM-HTN comorbidity population remains limited. Objective To methodologically explore and preliminarily evaluate an interpretable machine learning framework for 1-year MACE prediction in hospitalized patients with coexisting T2DM and HTN using routine clinical data. Methods This retrospective study included 1,054 hospitalized patients with T2DM and HTN, of whom 249 (23.6%) experienced MACE during 1-year follow-up. The dataset was randomly divided into training (60%), validation (20%), and independent test (20%) cohorts using stratified sampling. LASSO regression was applied for feature selection from 69 clinical variables. Four algorithms, including logistic regression, random forest, support vector machine, and XGBoost, were developed and compared. Model performance was assessed using discrimination, calibration, and clinical utility metrics. SHapley Additive exPlanations (SHAP) were used to interpret the final model. Results LASSO identified six stable predictors: HbA1c, age, hypertension duration, cystatin C (CysC), T2DM duration, and carotid intima-media thickness (CIMT). Sex was additionally incorporated based on clinical relevance. Multivariable logistic regression showed that HbA1c, age, hypertension duration, T2DM duration, CysC, and CIMT were associated with 1-year MACE risk, whereas sex was not statistically significant. Logistic regression showed the best relative balance between discrimination, calibration, and simplicity on the validation set, although learning curves indicated limited incremental improvement with increasing training sample size. After isotonic regression recalibration, the final logistic regression model achieved an ROC-AUC of 0.828, a PR-AUC of 0.656, and a Brier score of 0.116 on the independent test set. Decision curve analysis indicated potential clinical net benefit. SHAP linked model predictions to glycemic burden, aging, cumulative disease exposure, renal-related risk, and subclinical atherosclerosis. Conclusion An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN. CysC provided additional prognostic information beyond its conventional role as a renal filtration marker, although this association should be interpreted as prognostic rather than causal. External validation is required before the model can be considered for clinical decision support.
An interpretable RF-based model using routine clinical data effectively predicted MCI in elderly patients with T2DM and provided clinically intuitive explanations of risk drivers, and may support risk-stratified cognitive screening and individualized management.
Beibei Dong, Le Wang, Wen-Wen Shi et al.· Scientific Reports· 0 citations
Development and validate an interpretable machine learning model to predict the 1- to 3-year risk of cardiovascular events in breast cancer patients by integrating baseline and treatment variables and identified endocrine therapy, anemia management therapy, and history of cerebrovascular disease as the top three predictors.
Luxin Wang, Rui Yan, Xin-Yu Zhu et al.· Frontiers in Oncology· 0 citations
Abstract Background Patients undergoing dialysis are at an elevated risk of cardiovascular events. This study aimed to develop machine learning (ML) prediction models to identify risk factors for major adverse cardiovascular events (MACE) in dialysis patients. Materials and Methods This retrospective study included 203 patients undergoing dialysis with a median age of 45.0 years and 64.0% male. The participants were divided into training and test sets in a 7:3 ratio. LASSO regression selected characteristic variables from patients’general information, laboratory tests, and echocardiographic parameters (including global longitudinal strain [GLS]). Eight ML models were constructed,and SHAP analysis evaluated feature importance. Results The incidence of MACE (including myocardial infarction, unstable angina, heart failure, and cardiovascular death) in dialysis patients was 38.92%. The average follow-up period was 18 months. LASSO regression identified eight feature variables. Among the ML models, AdaBoost demonstrated superior performance, with an AUC of 0.883 (95% CI: 0.830–0.937), accuracy of 0.804, sensitivity of 0.864 and specificity of 0.762 in the training set, and an AUC of 0.809 (95% CI: 0.706–0.912), accuracy of 0.750, sensitivity of 0.90 and specificity of 0.675 in the test set. The SHAP analysis identified N-terminal pro-brain natriuretic peptide (NT-proBNP) level, estimated glomerular filtration rate (eGFR), GLS and age as the four most important features for predicting MACE in patients undergoing dialysis (mean absolute SHAP values: 0.199, 0.176, 0.096 and 0.091, respectively). Conclusion Elevated NT-proBNP, advanced age, reduced eGFR and impaired GLS were independently associated with an increased risk of MACE in patients undergoing dialysis.
Mei Jin, Zikang Lin, Lingxiang Ma et al.· Annals medicus· 0 citations
Patients with diabetic foot (DF) have a high risk of cardiovascular (CV) death, yet dedicated risk-prediction tools for this population are lacking. We developed and temporally validated an interpretable machine learning (ML) model for predicting CV death in patients with DF. This single-center retrospective cohort study included 2,835 patients admitted between February 2017 and May 2025. The development cohort comprised 2,325 patients, including 748 CV deaths, and was divided into training, internal validation, and held-out test sets; an independent temporal validation cohort included 510 patients, including 220 CV deaths. Nine supervised ML algorithms were compared using the area under the receiver operating characteristic curve (AUC). Extreme gradient boosting (XGBoost) showed the best overall performance. The optimal model, incorporating demographic and diabetes-related characteristics, routine laboratory parameters, and DF-specific features, achieved AUCs of 0.829 (95% confidence interval [CI]: 0.752-0.905) in internal validation, 0.844 (95% CI: 0.806-0.881) in the held-out test set, and 0.828 (95% CI: 0.789-0.868) in temporal validation. The model demonstrated good calibration and favorable net benefit on decision curve analysis. SHapley Additive exPlanations (SHAP) identified age, serum creatinine, glycated hemoglobin, triglycerides, and body mass index as the most influential predictors of increased model-predicted risk. This interpretable XGBoost model may support early identification and individualized risk stratification of patients with DF at high risk of CV death; however, prospective multicenter validation is required before clinical implementation.
Xiaoling Wan, Ting Shi, Qiao Liu et al.· Biomolecules & biomedicine· 0 citations
Background Older patients with type 2 diabetes mellitus (T2DM) and cardiovascular disease (CVD) frequently experience prolonged length of stay (PLOS). This condition increases healthcare burden and worsens prognosis. However, no predictive model specifically addresses PLOS in this high-risk multimorbid population. Methods This single-center retrospective study included hospitalized older T2DM-CVD patients. PLOS was defined as hospital stay exceeding the 75th percentile of the training set population. Potential predictors were selected via LASSO regression. Eight machine learning (ML) models were developed to predict PLOS risk. Model performance was evaluated using receiver operating characteristic curves, calibration curves, and decision curve analysis. SHAP analysis was employed for model interpretability. Results A total of 27,629 patients were included. The XGBoost model achieved the highest training AUC (0.819) and demonstrated competitive predictive performance in both the internal (AUC = 0.753) and time-based external (AUC = 0.728) validation sets. However, its performance advantage over simpler models such as logistic regression was modest in validation, and XGBoost showed some degree of overfitting (AUC drop of 0.066 from training to validation). Although logistic regression showed comparable validation performance with less overfitting, XGBoost was selected as the final model for its ability to capture complex nonlinear interactions and provide SHAP-based interpretability, with the understanding that further external validation is needed. Key predictors included cerebral infarction, white blood cell count, anemia, pulse rate, the glycated hemoglobin to high-density lipoprotein cholesterol ratio (GHR), and osteoporosis. Most continuous variables showed nonlinear associations with PLOS risk. Conclusions The XGBoost-based model effectively predicts PLOS risk in older T2DM-CVD patients. This tool shows promise for early identification of high-risk individuals and optimization of medical resource allocation within our institutional setting. However, further prospective and multi-center validation studies are required before clinical adoption.
Yixia Zuo, Jianfei Chen, Jie Wang et al.· Frontiers in Cardiovascular...· 0 citations
Hypertension is one of the most important modifiable risk factors for Cardiovascular Disease (CVD), yet identifying which hypertensive patients are at higher risk remains challenging in clinical practice. This study developed and evaluated three machine-learning models: logistic regression, random forest, and Gradient Boosting for CVD risk prediction in a cohort of 23,543 hypertensive patients drawn from a 70,000 patient cardiovascular dataset. After preprocessing, feature engineering, SMOTE-based class balancing, and hyperparameter tuning via randomized search, model performance was assessed on a held-out test set and validated using 5-fold stratified cross-validation with SMOTE correctly nested inside each fold to avoid data leakage. On the test set, tuned Gradient Boosting model achieved the highest accuracy (78.59%) and AUC-ROC (0.6681), outperforming Logistic Regression (0.6633) and Random Forest (0.6508). cross-validation provided a slightly different perspective: Logistic Regression’s mean AUC-ROC (0.6628) edged out Gradient Boosting (0,6609) and Random Forest (0.6383), SHAP analysis on the Gradient Boosting model identified systolic blood pressure, age, and height as the strongest predictors, with height rivaling systolic blood pressure and surpassing BMI a notable difference from Random Forest’s feature importance ranking. Lifestyle factors (smoking, alcohol, physical activity) contributed minimally. These findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.
C. M. Anyanwu, J. C. Onyianta, Ogechi Gift Onyedi et al.· Nature Journal of Emerging S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.