Data-driven fresh property prediction and mix design optimization of self-compacting geopolymer concrete
Abstract
Self-compacting geopolymer concrete (SCGC) requires coordinated control of workability and strength during mix design, yet predicting fresh-state properties from mix proportions alone has not been reliably achieved. This study trained machine learning models on 327 two-part SCGC mixes from 37 published sources to predict five fresh properties (slump flow, T500, V-funnel time, L-box ratio, J-ring step) and 28-day compressive strength (CS₂₈d). Curing temperature, pre-demoulding duration, and post-curing regime were added as inputs for CS₂₈d, giving 15 inputs total. Five tree-based algorithms (random forest, gradient boosting, extra trees, XGBoost, LightGBM) were compared, with hyperparameters tuned via RandomizedSearchCV or Optuna. Under random-split cross-validation, extra trees achieved CV R² of 0.949 (V-funnel), 0.910 (L-box), and 0.934 (J-ring); gradient boosting led for slump flow at CV R² = 0.929. Under Leave-One-Source-Out (LOSO) validation; which withholds entire laboratories from training; R² fell to 0.261 for slump flow and 0.120 for CS₂₈d; T500 reached R² = −0.279. The resulting gap, ΔR² = 0.67–0.92 across outputs, measures the inter-laboratory information leakage that random-split validation conceals and that prior SCGC ML studies have not corrected for. Adding the curing inputs raised CS₂₈d test R² by 0.119. SHAP and permutation importance analysis identified curing temperature as the dominant driver of CS₂₈d and produced physically consistent rankings for fresh property outputs. The trained models served as surrogates in a differential evolution framework that simultaneously optimises for EFNARC workability classes (SF2/SF3, PA2) and minimum compressive strength targets, tested for ambient (25°C, 24 h) and oven (70°C, 48 h) curing across six strength thresholds. Eleven of twelve scenarios returned feasible solutions, with total binder content from 362 to 574 kg/m³.