Exploring prognostic factors in breast cancer: development and selection of optimal machine learning models
Objective To investigate prognostic factors for breast cancer recurrence and metastasis, and to systematically compare multiple machine learning models to develop an optimal predictive tool. Methods We retrospectively analyzed data from 1,056 breast cancer patients diagnosed between January 2012 and June 2021 at a single center. Patients were randomly divided into a training (n=740) and a validation (n=316) set. Univariate and multivariate Cox proportional hazards regression analyses were performed to identify independent prognostic factors. Seven machine learning algorithms (Cox regression, LASSO, Elastic-Net, Decision Tree, Random Forest, XGBoost, and GBM) were employed. All models were implemented using survival-specific adaptations. A rigorous 5-fold cross-validation framework was used for model training and hyperparameter tuning. Model performance was evaluated using time-dependent Area Under the Curve (AUC), Brier scores, calibration curves, calibration-in-the-large, calibration slopes, and Decision Curve Analysis (DCA). SHAP values were employed for model interpretation. Results Multivariate Cox regression revealed that tumor size (cm) (HR = 1.025, 95%CI: 1.013-1.037), lymph node dissection (HR = 0.278, 95%CI: 0.199-0.389), ER% (HR = 1.006, 95%CI: 1.003-1.009), PR% (HR = 1.005, 95%CI: 1.002-1.008), Ki-67% (HR = 1.012, 95%CI: 1.007-1.016), and HER2 status (HR = 1.195, 95%CI: 1.098-1.301) were independently associated with disease-free survival. Random Forest and XGBoost demonstrated superior and stable predictive performance, with Random Forest achieving time-dependent AUCs of 0.851 (95%CI: 0.802-0.900, 1-year), 0.763 (95%CI: 0.702-0.824, 3-year), and 0.826 (95%CI: 0.766-0.886, 5-year) in the validation set. The Brier scores for Random Forest were consistently low, and calibration metrics confirmed excellent calibration. DCA indicated a positive net benefit across a wide range of threshold probabilities. Conclusion This single-center study identifies key prognostic factors for breast cancer and demonstrates that ensemble machine learning models, particularly Random Forest, offer superior predictive power. The integration of SHAP interpretation provides a methodological framework for potential clinical application. However, all predictive models require external, multi-center validation before clinical consideration. These findings provide a promising methodological basis, but caution is warranted against overinterpretation until independent verification is complete.