Reproducible and Explainable Machine Learning for Breast Cancer Classification: Sensitivity-Oriented Thresholding and Independent Methodological Replication
Abstract
Background: WDBC is a small historical benchmark, and near-ceiling performance alone provides limited evidence of transportability. Methods: We evaluated five model families for discrimination, calibration, and paired statistical testing. We utilized sensitivity-oriented out-of-fold (OOF) thresholds and decision curve analysis (DCA) and conducted an independent TOMPEI-CMMD methodological replication (larger two-center mammography cohort with biopsy-confirmed diagnoses; final cohort: 1358 patients, 1380 breasts; held-out: 272 patients, 279 breasts). Results: Across 20 additional stratified WDBC partitions, ROC-AUC remained consistently high (model means 0.988–0.995), but the criterion-specific nominal leader changed across partitions. Held-out TOMPEI ROC-AUC ranged from 0.775 to 0.804. No model demonstrated statistically significant superiority in ROC-AUC or frozen-threshold balanced accuracy after patient-cluster paired bootstrap comparison and Holm correction. The independent methodological replication yielded more moderate discrimination than the WDBC benchmark. Frozen sensitivity-oriented OOF thresholds transported reasonably but not uniformly. Conclusions: Multi-dimensional evaluation is more defensible than nominal AUC ranking. Neither the WDBC benchmark results nor the TOMPEI methodological replication establish population screening performance or clinical deployment readiness; further prospective clinically representative validation is required.