Skip to content
Open access

Predicting treatment-related cardiovascular risks in breast cancer patients: development and validation of an interpretable machine learning model

Aug 2026 · Frontiers in Oncology · Vol 16 · 0 citations · 26 references
Medicine

TL;DR

Development and validate an interpretable machine learning model to predict the 1- to 3-year risk of cardiovascular events in breast cancer patients by integrating baseline and treatment variables and identified endocrine therapy, anemia management therapy, and history of cerebrovascular disease as the top three predictors.

Abstract

Purpose This study aimed to develop and validate an interpretable machine learning model to predict the 1- to 3-year risk of cardiovascular events in breast cancer patients by integrating baseline and treatment variables, while preliminarily investigating the potential association between short-term cardiac function decline and long-term adverse cardiovascular events. Methods We analyzed electronic medical records from 31,878 breast cancer patients. A composite cardiovascular event outcome was used. Predictors were selected via a two-step process: removing highly correlated variables (|r|≥0.7) and applying LASSO regression with 10-fold cross-validation, which refined 62 initial variables down to 18. Five models were built and compared using the area under the receiver operating characteristic curve (AUC-ROC). The optimal model was interpreted using SHapley Additive exPlanations (SHAP). Results Among 31,878 breast cancer patients, 3,960 (12.4%) experienced cardiovascular events. The XGBoost model demonstrated the best overall discriminative performance (AUC = 0.790). SHAP analysis identified endocrine therapy, anemia management therapy, and history of cerebrovascular disease as the top three predictors. Crucially, short-term decline in cardiac function was also selected as a significant predictor, supporting its role as a precursor to long-term events. Model robustness was confirmed via sensitivity analysis.

Read PDF

Similar papers

Open access Jul 2026

Interpretable machine learning for cardiovascular disease risk prediction in cancer survivors: development and internal validation.

BACKGROUND Cardiovascular disease (CVD) is a major concern among cancer survivors. However, the intersection of cancer and CVD has only recently gained broader attention, and substantial evidence gaps remain. This study aimed to identify risk factors associated with incident CVD in cancer survivors and to develop a machine learning model for CVD risk prediction. METHODS In this retrospective study, we included 2,500 patients receiving systemic antitumor therapy at Jilin Cancer Hospital; 188 incident CVD events were observed. Variables spanning demographic, clinical, tumor- and treatment-related, and laboratory domains were collected. The dataset was randomly split into a training set (75%) and an internal validation set (25%). Feature selection was performed using LASSO regression. Six prediction models were developed, including five machine learning algorithms (GBM, CoxBoost, XGBoost, SVM, and RSF) and a traditional Cox proportional hazards model. Model performance was evaluated using time-dependent receiver operating characteristic curves (AUC), calibration analyses, and decision curve analysis. Model interpretability was assessed using SHapley Additive exPlanations (SHAP). RESULTS In univariate analyses, 48 variables were associated with CVD risk (P < 0.20). LASSO regression identified 19 predictors for model development. Key predictors included elevated systolic blood pressure, specific cancer types, anthracycline use, and a history of hypertension. The XGBoost model demonstrated the best predictive performance, with an average AUC of 0.666, sensitivity of 69.8%, specificity of 60.7%, and overall accuracy of 61.4% in the internal validation set. The model showed good calibration and yielded a positive net benefit across a range of clinical thresholds. SHAP analysis indicated that cancer type, lack of anti-HER2 therapy, elevated systolic blood pressure, advanced T stage, advanced TNM stage, and higher uric acid levels were the most influential predictors. CONCLUSION This study developed and internally validated interpretable machine learning models to predict CVD risk among cancer survivors. The models demonstrated good discrimination and calibration and outperformed traditional methods. By enabling individualized risk quantification and providing transparent interpretation of key predictors, this approach offers a practical tool to support personalized surveillance and prevention strategies in cardio-oncology.

Guoxing Zhang, Xueying Zhang, Chunyi Jia et al. · 0 citations
Open access Jul 2026

An interpretable machine learning model for predicting 1-year major adverse cardiovascular events in patients with type 2 diabetes and hypertension

An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN.

Juan Lv, Xi-Rui Wang, Zhengyi Zhang · 0 citations
Sep 2026

Improving 10-year cardiovascular disease risk prediction using automated machine learning.

AIMS To develop a cardiovascular disease (CVD) risk prediction model with improved accuracy and interpretability by integrating diverse risk factors and applying Automated Machine Learning (AutoML), thereby enhancing clinical utility over conventional models. METHODS This is a prospective cohort study. Data were obtained from the Multi-Ethnic Study of Atherosclerosis (MESA), including baseline and fifth follow-up visits, comprising 4713 participants. Exercise and dietary data were harmonized via Metabolic Equivalent of Task (MET) and Healthy Eating Index-2015 (HEI-2015), respectively. Predictor selection was performed using the Boruta algorithm alongside Random Forest (RF) error rate cross-validation. Logistic regression, four traditional machine learning algorithms, and H2O AutoML were each applied for model training and evaluation. Finally, the best-performing model was further interpreted using SHapley Additive exPlanations (SHAP). RESULTS A total of 21 predictors were selected, including age, sex, and Total Cholesterol (TC). Among the evaluated models, H2O AutoML outperformed other methods with an accuracy of 0.864, specificity of 0.892, precision of 0.610, F1 score of 0.670, and a Youden index of 0.635, achieving the highest AUC of 0.882 (0.846-0.918). SHAP analysis revealed the relative importance of predictors, with age, TC and Digit Symbol Score (DSS) ranking highest. CONCLUSIONS This study developed an AutoML-based CVD risk prediction model with superior discrimination and calibration, providing clinicians a practical tool for risk stratification. By enabling personalized prevention and early identification of high-risk individuals, this model has the potential to reduce CVD burden at the population level. Notably, DSS exhibited high importance and may represent a candidate risk marker.

Unknown authors · 0 citations
Open access Aug 2026

PREDICTING HEART DISEASE RISK FROM CLINICAL VARIABLES: A GENDER-SPECIFIC MACHINE LEARNING ANALYSIS AMONG HIGH-CHOLESTEROL PATIENTS

Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical decision support.

Taiwo Samson Adeyemo · 0 citations
Open access Aug 2026

Machine Learning- Based Cardiovascular Disease Risk Prediction in Hypertensive Patients: Explainable insights into Clinical Risk Factors

Hypertension is one of the most important modifiable risk factors for Cardiovascular Disease (CVD), yet identifying which hypertensive patients are at higher risk remains challenging in clinical practice. This study developed and evaluated three machine-learning models: logistic regression, random forest, and Gradient Boosting for CVD risk prediction in a cohort of 23,543 hypertensive patients drawn from a 70,000 patient cardiovascular dataset. After preprocessing, feature engineering, SMOTE-based class balancing, and hyperparameter tuning via randomized search, model performance was assessed on a held-out test set and validated using 5-fold stratified cross-validation with SMOTE correctly nested inside each fold to avoid data leakage. On the test set, tuned Gradient Boosting model achieved the highest accuracy (78.59%) and AUC-ROC (0.6681), outperforming Logistic Regression (0.6633) and Random Forest (0.6508). cross-validation provided a slightly different perspective: Logistic Regression’s mean AUC-ROC (0.6628) edged out Gradient Boosting (0,6609) and Random Forest (0.6383), SHAP analysis on the Gradient Boosting model identified systolic blood pressure, age, and height as the strongest predictors, with height rivaling systolic blood pressure and surpassing BMI a notable difference from Random Forest’s feature importance ranking. Lifestyle factors (smoking, alcohol, physical activity) contributed minimally. These findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.

C. M. Anyanwu, J. C. Onyianta, Ogechi Gift Onyedi et al. · 0 citations
Open access Jul 2026

Exploring prognostic factors in breast cancer: development and selection of optimal machine learning models

Objective To investigate prognostic factors for breast cancer recurrence and metastasis, and to systematically compare multiple machine learning models to develop an optimal predictive tool. Methods We retrospectively analyzed data from 1,056 breast cancer patients diagnosed between January 2012 and June 2021 at a single center. Patients were randomly divided into a training (n=740) and a validation (n=316) set. Univariate and multivariate Cox proportional hazards regression analyses were performed to identify independent prognostic factors. Seven machine learning algorithms (Cox regression, LASSO, Elastic-Net, Decision Tree, Random Forest, XGBoost, and GBM) were employed. All models were implemented using survival-specific adaptations. A rigorous 5-fold cross-validation framework was used for model training and hyperparameter tuning. Model performance was evaluated using time-dependent Area Under the Curve (AUC), Brier scores, calibration curves, calibration-in-the-large, calibration slopes, and Decision Curve Analysis (DCA). SHAP values were employed for model interpretation. Results Multivariate Cox regression revealed that tumor size (cm) (HR = 1.025, 95%CI: 1.013-1.037), lymph node dissection (HR = 0.278, 95%CI: 0.199-0.389), ER% (HR = 1.006, 95%CI: 1.003-1.009), PR% (HR = 1.005, 95%CI: 1.002-1.008), Ki-67% (HR = 1.012, 95%CI: 1.007-1.016), and HER2 status (HR = 1.195, 95%CI: 1.098-1.301) were independently associated with disease-free survival. Random Forest and XGBoost demonstrated superior and stable predictive performance, with Random Forest achieving time-dependent AUCs of 0.851 (95%CI: 0.802-0.900, 1-year), 0.763 (95%CI: 0.702-0.824, 3-year), and 0.826 (95%CI: 0.766-0.886, 5-year) in the validation set. The Brier scores for Random Forest were consistently low, and calibration metrics confirmed excellent calibration. DCA indicated a positive net benefit across a wide range of threshold probabilities. Conclusion This single-center study identifies key prognostic factors for breast cancer and demonstrates that ensemble machine learning models, particularly Random Forest, offer superior predictive power. The integration of SHAP interpretation provides a methodological framework for potential clinical application. However, all predictive models require external, multi-center validation before clinical consideration. These findings provide a promising methodological basis, but caution is warranted against overinterpretation until independent verification is complete.

Meiying Shen, Yulei Wang, Zong-ming Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.