Skip to content

Predictive performance investigation of reactive extraction of caproic acid using explainable machine learning models

Nov 2026 · Chemical Engineering Communications · Vol 213, pp. 2061 - 2071 · 0 citations · 43 references

TL;DR

The present study explores the application of machine learning techniques to predict the E% for caproic acid systems using available variables prior to experimentation, and the XGBoost model shows the best performance.

View source

Similar papers

Open access Oct 2026

Machine learning-based prediction of chlorophenol removal from wastewater using reverse osmosis

One of the most popular methods for reducing the presence of extremely toxic compounds in wastewater is the reverse osmosis (RO) process. The purpose of this research was to examine the prosperity of different machine learning (ML) techniques based optimisation for the removal of chlorophenol from wastewater using a single module of RO process. The main intention was to predict the optimal operating conditions that would introduce a maximum chlorophenol rejection using different ML techniques (the Linear Regression, Ridge, Lasso, Decision Tree, Random Forest, Gradient Boosting, SVR, XGBoost, and Neural Networks) based on available experimental data. The performance of the developed models was evaluated using root mean square error (RMSE), mean absolute error (MAE), coefficient of determination (R²), and cross-validation. The results demonstrated that the Gradient Boosting algorithm can achieve the best performance in predicting chlorophenol removal, followed by the XGBoost algorithm. However, neural networks performed the worst. The results show that feed concentration and pressure are the essential operating conditions that can dominate the rejection rate of chlorophenol. More importantly, it was concluded that feed pressure of 12.76 atm, feed concentration of 5079.05 kmol/m³, wastewater temperature of 31.11 °C, and feed flowrate of 2.36 × 10 -4 m³/s are the optimal conditions to attain the maximum chlorophenol rejection of 82.37%.

E. M. Hameed, Ahmed Abdul Azeez Ismael, Ali Mahmood Khalaf · 0 citations
Open access Aug 2026

Revisiting the LSER Approach in the Era of Machine Learning: Insights from IAM Chromatography

The present study demonstrates the integration of the linear solvation energy relationship (LSER) concept with machine learning (ML) methodologies to improve the predictive and interpretative capabilities of chromatographic retention modeling. Immobilized artificial membrane (IAM) chromatography was employed as a model biochromatographic system, and an in-house library of 993 structurally diverse compounds, with experimentally determined chromatographic hydrophobicity index of IAM (CHIIAM), was used to train LSER-ML models. LSER descriptors were calculated using Absolv and extended with ionization-state descriptors to evaluate the applicability of the LSER framework for realistic in silico virtual screening scenarios. Several regression algorithms were tested, including linear, neighborhood-based, kernel-based, and ensemble tree-based models. Among them, the support vector regression with the radial basis function kernel (SVR-RBF) demonstrated the most balanced performance across all validation metrics of R2train = 0.884, R2test = 0.853, and Q2cv = 0.811, achieving predictive errors (RMSEtrain = 5.546, RMSEtest = 4.609, and RMSEcv = 6.989) close to the analytical uncertainty. Model interpretability was achieved using SHapley Additive exPlanations (SHAP), which confirmed the mechanistic relevance of the Abraham descriptors and the dominant contribution of hydrophobic volume and hydrogen-bonding properties to IAM retention. Applicability domain was verified with a Williams plot (±3 standardized residuals and leverage threshold h*). The results indicate that the proposed LSER-ML approach provides an interpretable, robust, and generalizable tool for modeling membrane-mimetic chromatographic systems and can be effectively applied in virtual screening and property-based molecular design.

W. Nisterenko, K. Greber, Magdalena Kierkowicz et al. · 0 citations
Open access Aug 2026

Flowability prediction and artificial intelligence analysis of SCC based on machine learning

Six machine learning algorithms were employed to construct artificial intelligence models for the precise prediction of self-compacting concrete (SCC) flow properties, and the extreme gradient boosting (XGB) model was identified as exhibiting superior predictive accuracy and generalization performance.

Jinlei Mu, Xin Fang, Ke-qiang Cao et al. · 0 citations
Open access Aug 2026

Machine Learning for Alkali-Activated Concrete: Feature Attribution, Strength–Carbon Relationships, and the Limits of Out-of-Campaign Generalisation

Machine learning (ML) models for alkali-activated concrete (AAC) are almost universally evaluated with random train–test splits, yet the literature-compiled datasets are strongly clustered by source study, and the reliability of such evaluations has rarely been quantified. The novelty of this study is a systematic quantification of out-of-campaign generalisation—via Leave-One-Study-Out (LOSO) cross-validation—for ML models trained on the largest curated public AAC dataset (1630 mixtures compiled from 106 published sources), together with model interpretation and an exploratory strength–carbon analysis. Four models (Linear Regression, Random Forest, Gradient Boosting, and optimised extreme gradient boosting, XGBoost) were benchmarked for predicting 28-day compressive strength (CS28). XGBoost performed best under conventional random splitting, with test-set coefficient of determination R2 = 0.801 and root-mean-square error (RMSE) = 7.21 MPa (5-fold cross-validation R2 = 0.758 ± 0.050). Under LOSO validation across 85 study folds, however, the median R2 collapsed to −0.328, with 49 of 85 folds negative: random-split metrics on literature-compiled AAC datasets are substantially inflated by within-study clustering, and study-stratified evaluation should become standard practice in this field. Within these limits, SHapley Additive exPlanations (SHAP) identified ground granulated blast-furnace slag (GGBFS) content, specimen geometry, CaO fraction, curing time, and sodium silicate (Na2SiO3) content as the five most influential predictors; because the oxide descriptors are derived from the declared binder proportions and the carbon-footprint values are inherited estimates from the source dataset, these attributions are associational rather than causal. No practically meaningful overall linear association was observed between estimated CO2 footprint and CS28 (Pearson r = −0.113, 95% CI [−0.175, −0.050], R2 = 0.013), and a Pareto analysis identified 14 candidate low-carbon, high-strength formulations for further experimental and life-cycle assessment. The developed models are suitable for within-dataset feature attribution and exploratory screening restricted to the represented feature domain; they should not be used as external mix-design tools without validation on independent experimental campaigns.

F. Pacheco-Torgal, Saqib Iqbal · 0 citations
Open access Aug 2026

Predicting Cellulose Concentration in Lyocell Slurry Using Hybrid Ensemble of Machine Learning Models

Accurate prediction of cellulose concentration in lyocell slurry is crucial for process control and product quality in sustainable lyocell fiber production, yet the complex, nonlinear nature of the swelling process makes it challenging to model using conventional parametric methods. This study develops a hybrid ensemble machine learning approach to predict the cellulose concentration of lyocell slurry based on industrial production data. A data set comprising 350 samples was collected from an industrial pulping process, covering 18 input variables related to raw material properties and process conditions. Six conventional machine learning modelsGaussian process regression (GPR), support vector regression (SVR), kernel regression (KR), multivariate adaptive regression spline (MARS), random forest (RF), and artificial neural network (ANN)were first established and optimized using Bayesian optimization with 5-fold cross-validation. Subsequently, a hybrid ensemble model (HEM) was constructed by aggregating the predictions of the six base models using a random forest meta-learner selected through score analysis. The predictive performance of all models was evaluated using multiple metrics, including coefficient of determination (R 2), root-mean-square error (RMSE), and mean absolute error (MAE). The results show that the HEM achieves the best overall testing performance (R 2 = 0.761, RMSE = 0.067, MAE = 0.053), followed closely by the random forest model (R 2 = 0.754, RMSE = 0.068, MAE = 0.053). The ANN exhibits the smallest training−testing performance gap, confirming the effectiveness of L2 regularization. Kernel-based methods and MARS yield inferior accuracy (testing R 2 < 0.70), indicating the limitations of global smoothness or additive assumptions for this task. SHAP analysis identifies the NMMO-to-cellulose ratio, hydroxylamine concentration, and temperature parameters as the most influential features. The proposed HEM provides a reliable tool for online prediction of slurry cellulose concentration, enabling real-time process control and contributing to more efficient and sustainable lyocell fiber production.

Huinan Yu, Meng Yan, Daqian Yu et al. · 0 citations

Related blog posts

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.