Source-Study-Aware Validation and Explainable Machine Learning for Reliable Shear Capacity Assessment of Steel-Fiber-Reinforced Concrete Corbels
This study proposes a source-study-aware validation framework to estimate the shear capacity of stirrup-free steel-fiber-reinforced concrete corbels and assess the reliability and generalizability of machine-learning models in structural engineering. A database of 108 specimens, compiled from seven independent studies, was analyzed using linear regression, ridge regression, and hyperparameter-optimized XGBoost. The evaluation combined standard specimen-level five-fold cross-validation, predefined source-study holdout evaluation, exhaustive partition-sensitivity analysis, and specimen-level out-of-bag bootstrap uncertainty analysis. In the predefined source-study hold-out evaluation, with hyperparameters selected exclusively from the training portion via an inner group-aware search, the models achieved test R2 values of 0.842, 0.866, and 0.912, respectively, with XGBoost providing the highest point estimate of performance. However, an exhaustive sensitivity analysis of source-study partitions, in which hyperparameters were re-selected independently within the training portion of each partition, indicated that this ranking was not maintained: Linear Regression achieved the highest mean and median R2 and the narrowest interquartile range, whereas XGBoost exhibited the lowest mean and median R2, the widest interquartile range, and the highest incidence of negative R2 values and large prediction errors among the three models; this reflects instability introduced by re-tuning under limited, group-diverse training data. The specimen-level out-of-bag bootstrap uncertainty analysis also favored the linear models in terms of mean performance and confidence-interval width. Standardized coefficients and SHAP analyses consistently identified the shear span-to-effective-depth ratio, the longitudinal reinforcement ratio, and the fiber ratio as the dominant predictors. This study provides a transparent and uncertainty-aware framework for determining when predictive performance is transferable across independent experimental studies and when caution is required in engineering applications.