Skip to content

Impact of Training Data Proportion on Machine Learning–Based Prediction of Soaked CBR

2026 · Journal of Structural Design and Construction Practice · 0 citations · 46 references

Abstract

The soaked California bearing ratio (CBR) is an essential parameter of transportation engineering to design flexible pavement. The laboratory determination of CBRs is a lengthy, time-consuming process. To assess the CBRs of pavement materials, this investigation introduces a robust artificial intelligence model by comparing the multilinear regression (MLR), linear (LSSVM_L), and polynomial (LSSVM_P) kernel-based least squares support vector machine (LSSVM) models. Moreover, the six training datasets were developed to determine the impact of the quality and quantity of the training database on the MLR and LSSVM models. The performance analysis reveals that the linear kernel-based LSSVM (referred to as LSSVM_L50) model is simple yet highly accurate, with reduced complexity, and achieves higher performance (0.9992 and 0.9602 in the training and testing phases, respectively) using a 50% training database. Still, the LSSVM_P70 model achieved CBRs performance of 0.9999 (in training) and 0.9791 (in testing), outperforming the LSSVM_L50 model, as the polynomial kernel effectively handles the database’s nonlinearity and complexity. In addition, the generalizability analysis, uncertainty analysis, and validation using external and laboratory-tested samples confirmed the robustness of the LSSVM_P70 model. Conversely, the Shapley additive explanations analysis revealed that the optimum moisture content dominates model predictions, with a mean absolute impact of 7.14, followed by gravel and fine content as secondary drivers, while demonstrating clear nonlinear relationships, threshold effects, and instance-level interpretability.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.