Physics-Informed Hybrid Machine Learning Model for Carbonation Depth Prediction in Concrete through Residual Correction and Variogram Analysis of Response Surfaces
Nov 2026· Journal of materials in civil engineering· 0 citations· 65 references
TL;DR
A hybrid residual correction framework that integrates a physics-based carbonation model with a stacked ensemble of machine learning algorithms: gradient boosted regression trees (GBRT), support vector regression (SVR), and Gaussian process regression (GPR), combined through an XGBoost metamodel, demonstrating that the residual-based metamodel reproduced observed carbonation depths with higher accuracy.
Abstract
Carbonation-induced deterioration of reinforced concrete is a major durability concern, as it reduces pore solution alkalinity and accelerates reinforcement corrosion. Conventional service life models often oversimplify the combined effects of material and environmental factors, limiting their predictive reliability. This study presents a hybrid residual correction framework that integrates a physics-based carbonation model with a stacked ensemble of machine learning algorithms: gradient boosted regression trees (GBRT), support vector regression (SVR), and Gaussian process regression (GPR), combined through an XGBoost metamodel. Unlike conventional stacking, the metamodel is trained on residuals between physical model predictions and experimental measurements. This enables systematic correction of mechanistic biases while retaining physical interpretability. A second contribution is the application of variogram analysis of response surfaces (VARS) for variance-based global sensitivity analysis, which quantifies feature influence using geostatistical indicators (sill, nugget, and range), offering insights beyond standard feature importance methods. The model was deployed as an app within a MATLAB-based user interface (UI) to promote practical use, enabling service life prediction from minimal, easily measurable input parameters. The framework was validated against accelerated carbonation experiments and long-term natural exposure data from the literature. The results demonstrate that the residual-based metamodel reproduced observed carbonation depths with higher accuracy than individual base learners or the physical model.
Developing reliable computational tools for durability and service-life assessment of concrete structures in aggressive environments is essential for advancing predictive modeling in structural engineering. This study introduces machine learning (ML)–based models for forecasting the sulfate and acid resistance of recycled aggregate geopolymer concrete (RGPC), produced with untreated and surface-treated recycled concrete aggregates (RCAs) through two mixing approaches. Three algorithms, i.e. Gaussian process regression (GPR), LSBoost ensemble, and Neural Network, were trained using nine input parameters related to material composition and exposure conditions, with durability indicators, namely mass loss rate (Kw) and compressive strength retention index (Kf), as outputs. A dataset of 336 experimentally tested RGPC specimens was used, applying Bayesian Optimisation for hyperparameter tuning and 5-fold cross-validation for generalisation. Among the models, the optimized GPR achieved the highest accuracy, confirmed by the lowest objective value. Feature importance analysis highlighted environmental cations, sulfate concentration, RCA replacement level, initial compressive strength, and exposure duration as the most influential factors governing degradation. The proposed Bayesian-optimized ML framework demonstrates a robust and generalizable method for predicting durability and service life of sustainable concretes, providing a valuable tool for simulation-driven design and durability-based performance assessment in mechanics and structural engineering.
P. Singh, Puja Rajhans· Engineering Research Express· 0 citations
Accurate estimation of bond strength between steel reinforcement and geopolymer concrete is essential for the reliable design of sustainable reinforced concrete structures. However, the highly nonlinear interactions reduce the applicability and accuracy of conventional empirical models. This study proposes a Bayesian-optimized interpretable machine learning framework to predict the ultimate bond strength of reinforced geopolymer concrete using a comprehensive experimental database compiled from published studies. A dataset of 238 samples with 20 influential input variables was assembled to represent material properties, geopolymer chemistry, and specimen geometry. Six advanced machine learning algorithms, including Support Vector Regression (SVR), Random Forest (RF), Extra Trees Regressor (ETR), Gradient Boosting Machine (GBM), XGBoost, and CatBoost, were developed and systematically compared. Hyperparameter tuning was performed using Bayesian optimization to improve model performance. The results indicate that all models achieved strong predictive capability, while the optimized CatBoost model (BO-CatBoost) provided the best performance with testing metrics of R² = 0.950, MAE = 1.173, MAPE = 11.608%, and RMSE = 1.669. A comparative evaluation with existing empirical equations further demonstrated the superior accuracy and lower prediction variability of the proposed model. To enhance model transparency, SHAP-based explainability analysis was conducted to quantify the contribution of each input parameter. The global importance analysis revealed that compressive strength, the embedment length-to-bar diameter ratio, and the cover-to-bar diameter ratio are the most influential factors governing bond strength. Additional mixture-related parameters, including the alkaline solution-to-binder ratio, curing temperature, CaO content in the binder, and the SiO₂/Al₂O₃ ratio, also contribute to the bond mechanism by influencing geopolymerization and matrix densification. The proposed framework provides both high predictive accuracy and interpretable insights, demonstrating the potential of Bayesian-optimized interpretable machine learning to support the design and optimization of sustainable reinforced geopolymer concrete structures.
Chloride-induced corrosion is one of the principal causes of deterioration in reinforced concrete infrastructure, making accurate prediction of chloride diffusion coefficients essential for durability assessment and service-life design. Existing machine learning models often suffer from limited experimental datasets and insufficient incorporation of engineering knowledge, restricting their predictive capability and generalization. This study presents a physics-guided machine learning framework that integrates domain-informed feature engineering, conditional synthetic data augmentation, and stacking ensemble learning to predict the chloride diffusion coefficient of concrete from Rapid Chloride Migration (RCM) test data. Physics-guided features were developed to represent fundamental transport mechanisms and binder characteristics, while synthetic data augmentation was employed to improve data coverage and enhance model robustness. The final stacking ensemble combined CatBoost, XGBoost, Random Forest, and Linear Regression through a Ridge Regression meta-learner. The proposed framework achieved a coefficient of determination (R2) of 0.903, with an RMSE of 1.321 and an MAE of 0.920 on an independent holdout dataset, outperforming all individual machine learning models. Ablation analysis demonstrated that synthetic data augmentation was the primary contributor to performance improvement, while ensemble learning provided additional gains in predictive accuracy and robustness. Model interpretability using SHapley Additive exPlanations (SHAP) identified slag content, water-to-binder ratio, and porosity-related variables as the dominant factors governing chloride diffusion predictions, consistent with established durability mechanisms. The proposed framework provides an accurate and interpretable tool for chloride diffusion prediction that supports durability assessment, service-life estimation, and the design of sustainable concrete mixtures.
Moutaman M. Abbas· Journal of Composites Scienc...· 0 citations
The significance of rock brittleness is well‐recognized in the fields of geotechnical engineering and energy exploration. To enhance the predictive precision of rock brittleness, this paper proposes a Stacking integrated algorithm. This algorithm synergistically combines various meta models and foundational models, utilizing a suite of nine algorithm modes: Gaussian Process Regression, Support Vector Machine, Backpropagation Neural network, Extreme Learning Machine network, Decision Tree, Random Forest, Extreme Gradient Boosting, Lasso Regression, and Ridge Regression. Furthermore, Tuna and Bayesian optimization algorithms are utilized to refine the model's performance. Additionally, a new diversity index, k, based on the ratio of correlation coefficients, has been introduced to facilitate the optimal selection of base models for the Stacking integrated algorithm. The predictive accuracy of the Stacking integrated model, as determined by the proposed diversity index k, surpasses that of the finest individual base model and outperforms other sets of five integrated base models with an equivalent number of components. This underscores the efficacy of the diversity index k in guiding the selection of appropriate base models for the stacking process. The most effective model for predicting rock brittleness incorporates the Extreme Learning Machine network, Decision Tree, Random Forest and Extreme Gradient Boosting. This ensemble model demonstrates superior accuracy over the single best model, the Decision Tree, by reducing the average brittleness prediction error rate by 0.3745 and elevating the determination coefficient (
R
2
) value from 0.9118 to 0.9629.When compared with the Particle Swarm Optimization model, this composite model achieves an increase of 0.1498 in
R
2
.
Jing Jia, Diquan Li, Ziyi Zhang et al.· International journal for nu...· 0 citations
Understanding permeability is essential for evaluating reservoir quality and field development planning. Reliable permeability estimation can reduce the uncertainty in reservoir characterization, particularly in intervals where core data are limited. As the industry relies on log-based interpretations and empirical correlations, the limitations of these approaches become apparent. Data-driven approaches offer a promising alternative to conventional empirical methods. The data set in this study comprises 252 samples with seven features derived from conventional well logs. Data preprocessing includes handling missing values, smoothing logs, feature engineering to add an extra input, and transformation with the Yeo-Johnson technique. A center moving average filter was used to reduce variance and improve data consistency. Ensemble machine-learning (ML) and baseline models were developed and evaluated using a 75-25 train-test split, four-fold cross validation, and model complexity assessment. Ensemble methods outperformed baseline models, with extremely randomized trees (ET) and random forest (RF) emerging as the most stable, achieving a mean R² of 0.93 and 0.89 and a low R² standard deviation (0.3). Multilinear regression (MLR) and artificial neural networks (ANNs) show limited accuracy, while gradient boosting (GB) and extreme gradient boosting (XGBoost) methods exhibit overfitting despite perfect training scores. Predicted kh values were compared with core data. Both linear and nonlinear empirical equations were derived using MLR, polynomial regression, and a power-law model. The power-law model (empirical equation) achieved an R² value of 0.79 and can therefore be used to estimate permeability. Additionally, a Gaussian mixture model (GMM) was used for unsupervised classification of hydraulic flow units (HFU) using the flow zone indicator (FZI), computed from the continuous permeability curve obtained from the best ML model. Thus, ML-based permeability prediction is an indispensable component of HFU modeling. The model identified three distinct flow zones, enabling HFU clustering and defining their corresponding petrophysical properties and depositional environments.
Vikram Kumar, Sayantan Ghosh, S. Maiti· Petrophysics· 0 citations
Accurate prediction of porosity and permeability in complex carbonate reservoirs is very important for understanding reservoirs, but remains challenging due to inherent heterogeneity. This study develops a robust, machine learning-driven workflow to enhance the prediction of these critical petrophysical properties and the identification of Hydraulic Flow Units. The methodology integrates conventional core data and geophysical well logs, employing advanced data preprocessing, including depth matching, which significantly improved the log-core porosity correlation. A key innovation involves using a Gaussian Mixture Model for unsupervised Hydraulic Flow Unit identification, which outperformed traditional empirical methods and K-Means clustering by yielding five distinct Hydraulic Flow Units with high intra-unit porosity-permeability correlations (R2 up to 0.93) validated by Mercury Injection Capillary Pressure data. For predictive modeling, a comprehensive comparison of algorithms revealed that a Voting ensemble meta-algorithm with a Multi-Layer Perceptron base learner delivered superior performance for both porosity (on integrated data from three wells) and permeability (modeled per Hydraulic Flow Unit). The final models successfully estimated properties in non-cored intervals and a blind well, demonstrating high accuracy and generalizability. This integrated approach provides a reliable and theory-grounded framework for characterizing heterogeneous carbonate reservoirs, reducing dependency on extensive coring operations.
Mostafa Khazaei Panah, Razieh Khosravi, M. Simjoo et al.· Scientific Reports· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.