Skip to content

Predicting soil-water partition coefficients of PFAS using machine learning: Model development, interpretation, and validation.

Jul 2026 · Environmental Pollution · pp. 128771 · 0 citations · 50 references
Medicine

Abstract

PER: and polyfluoroalkyl substances (PFAS) are environmentally persistent contaminants, yet experimental determination of their soil-water partition coefficients (Kd) remains costly and time-consuming. In this study, five machine-learning regression models were developed using 2057 literature-derived batch adsorption data points by integrating the average net charge descriptor (Zavg), the composite descriptor alert_prior_score, equilibrium aqueous concentration (Cw), PFAS structural descriptors, and soil physicochemical properties. Among the tested models, extreme gradient boosting (XGBoost) showed the best performance with 10 descriptors, achieving R2 values of 0.83 and 0.86 and ratio of performance to deviation (RPD) values of 2.42 and 2.66 for the Cw < 10 μg/L and Cw ≥ 10 μg/L datasets, respectively. Shapley additive explanations (SHAP) analysis indicated that hydrophobic interactions dominated adsorption at low concentrations (Cw < 10 μg/L), whereas headgroup-related hydrophilicity became more influential at higher concentrations. Independent sorption experiment validation using contaminated site soils showed that prediction deviations for all samples were within one order of magnitude. These results demonstrate that the proposed model provides an efficient and interpretable tool for predicting PFAS soil-water partitioning and supports environmental risk assessment and contaminated-site management.

View source

Similar papers

Aug 2026

Interpretable machine learning for predicting gaseous arsenic adsorption by metal oxides and identifying influential descriptors.

Identifying descriptors associated with gaseous arsenic adsorption by metal oxides remains challenging because literature data are heterogeneous and incomplete. A database of 280 experimental records and 20 descriptors from 17 studies was compiled to predict adsorption capacity and interpret descriptor-performance relationships. Mean, k-nearest neighbor (KNN), and inference-based imputation strategies were combined with gradient boosting decision tree (GBDT) and particle swarm optimization-tuned GBDT (PSO-GBDT) models. Among six configurations, the PSO-GBDT model trained on the inference-imputed dataset achieved the lowest five-fold cross-validation RMSE of 2.10 mg/g and test-set R2, RMSE, and MAE values of 0.98, 1.54 mg/g, and 0.71 mg/g, respectively. Permutation feature importance (PFI) and Shapley additive explanations (SHAP) showed that operating and gas-phase descriptors dominated predictions, with H2O concentration and adsorption time ranked highest, followed by adsorption temperature, As2O3 concentration, Fe content, and average pore diameter. Partial dependence plots (PDPs) associated higher predicted capacities with longer adsorption times, higher As2O3 concentrations, larger pore diameters, and lower adsorption temperatures within the compiled data range. As an exploratory application, the model prioritized Fe-Mn adsorbents containing 63%-81.5% Fe and 18.5%-37% Mn with pore diameters of 20-24 nm, corresponding to predicted capacities of 20-21.2 mg/g. Subsequent screening over broader operating ranges predicted a high-capacity region of 53-54.1 mg/g at As2O3 concentrations of 110-200 ppm, adsorption times of 48-240 min, and temperatures of 300-600 °C. Overall, the framework supports interpretable prediction and hypothesis generation for gaseous arsenic adsorption by metal oxides.

Yanhong Zhu, Qi Liu, Shuang-chun Wen et al. · 0 citations
Open access Sep 2026

Research on a Prediction Model for the Bioconcentration Factor (BCF) of Polyhalogenated Organic Phosphates Based on QSPR

The bioconcentration factor (BCF) is central to ecological-risk assessment, but experimental BCF measurement is too slow for large-scale chemical screening. Polyhalogenated organophosphate esters are widely used flame retardants and remain of concern because of their persistence and potential bioaccumulation. Here, we developed a quantitative structure–property relationship (QSPR) framework to predict BCF and support environmental risk prioritization for structurally related compounds. A dataset of 160 compounds was divided into training (n = 130) and test (n = 30) sets. Ten descriptors were selected from 766 candidates using a genetic algorithm. Multiple linear regression (MLR), support vector machine (SVM), and backpropagation artificial neural network (BP-ANN) models were constructed and evaluated. Dataset partitioning was examined using Tanimoto similarity analysis and uniform manifold approximation and projection (UMAP) visualization. Model performance was assessed using 11 validation metrics, bootstrap confidence intervals, and Williams-plot applicability-domain analysis. Among the three models, BP-ANN gave the lowest external prediction errors in the present test set (MAEtest = 0.33 and RMSEtest = 0.41) and Q2F2 = 0.79. This ranking should be interpreted cautiously because the test set contained 30 compounds. Descriptor interpretation suggests that BCF variation is associated with lipophilicity, molecular topology, electronic distribution, and phosphorus-containing functional groups. The framework may support early-tier screening of structurally related flame retardants within the defined applicability domain, but experimental confirmation remains necessary for regulatory decisions.

Xiongjun Yuan, Cheng Wang, Yong-De Wei et al. · 0 citations
Open access Jul 2026

Assessing Features of PFAS Groundwater Occurrence Using SHAP-Enhanced Machine Learning.

Per- and polyfluoroalkyl substances (PFAS) are persistent groundwater contaminants that pose long-term risks to water sources and public health. Predicting PFAS occurrence remains challenging due to high-dimensional environmental data and limited model interpretability. In this study, we benchmarked explainable machine learning models to predict PFAS occurrence in groundwater and to identify the features strongly associated with estimated PFAS occurrence. A comprehensive dataset of 12,406 groundwater characterization records collected between 2001 and 2019 was compiled with 172 explanatory features to describe PFAS source proximity, land use, hydrogeology, soil properties, meteorology, and sampling sites characteristics. Four tree-based ensemble classifiers, including Random Forest, XGBoost, LightGBM, and CatBoost, were evaluated under multiple classification schemes. Binary classification, which split total PFAS concentrations into 'low' and 'high' categories based on the median value, achieved the most robust and generalizable performance, with testing accuracy above 94% and well-established precision-recall curves. Increasing the number of classification bins degraded performance, particularly for intermediate bins, highlighting intrinsic separability limits in PFAS occurrence data rather than model deficiencies. Model interpretability was addressed using SHapley Additive exPlanations (SHAP), which revealed that sampling year and proximity to major PFAS sources were the dominant predictors across all models. Additional contributions were attributed to proximity to other PFAS sources, soil texture, hydrologic features, land use, and precipitation patterns. SHAP interaction analyses further revealed model-learned temporal variation in the attribution of source-proximity predictors. These patterns are interpreted as hypothesis generating model associations that may be related to regulatory changes, evolving monitoring strategies, analytical-era differences, and possible secondary-source influences, rather than as confirmation of specific environmental processes. Collectively, this study achieved interpretable prediction for PFAS occurrence in groundwater, and offered a transparent, data-driven framework to inform PFAS risk assessment.

Lin Wang, Yun Ma, D. Rajapakshe et al. · 0 citations
Open access Jul 2026

Machine Learning Prediction and Interpretation of Soil−Water Characteristic Curves of Biochar-Amended Soils

Biochar is a porous, carbon-rich soil amendment that can enhance soil water retention capacity by modifying pore structure and physicochemical properties. Understanding the soil−water characteristic curve (SWCC) of biochar-amended soils is essential for evaluating their hydrological behavior and promoting the application of biochar in engineering practice. Given the demonstrated feasibility and accuracy of machine learning methods for predicting soil parameters, this study employed six machine learning models, namely, decision tree, random forest, XGBoost, LightGBM, CatBoost, and artificial neural network, to predict the SWCC of biochar-amended soils based on a constructed dataset. Feature importance analysis and partial dependence analysis were further conducted to reveal the influence patterns of key variables. The results indicate that all six models exhibit good predictive capability, with gradient boosting models (XGBoost, CatBoost, and LightGBM) performing best. Suction is the dominant factor controlling the volumetric water content variation, while soil particle-size distribution and dry density provide the physical basis for water retention. Biochar content, pyrolysis temperature, and feedstock type further modulate the water retention capacity of amended soils. Overall, the findings demonstrate that machine learning approaches can effectively predict the SWCC of biochar-amended soils and provide insights into the controlling mechanisms of soil water retention.

Yu Luo, Letian Wang, Zixuan Zheng et al. · 0 citations
Open access Jul 2026

Water remediation using sustainable kaolin-derived zeolite and machine learning-guided prediction of adsorption performance.

A chemistry-informed machine learning (CIML) model was developed for the prediction of equilibrium concentration (Ce) and adsorption capacity (Qe) of synthesized sustainable and cost-effective zeolite 4 A from kaolin clay. In this framework, the chemistry-informed integrates fundamental adsorption principles, including mass balance, kinetic descriptors (mass transfer dynamics), equilibrium relationship (saturation behaviour), stoichiometry consistency and dimensionless loading, into a learning architecture. The materials are thoroughly characterized using FT-IR, FE-SEM, XRD, and BET. In the experimental study, 1 g of zeolite was utilized to treat 10 L of dye effluent through column adsorption demonstrating the material's practical applicability for large volume wastewater treatment. The breakthrough curve exhibited sigmoidal profile, with Thomas model yielding an adsorption capacity of 56.6 mg/g. Extensive experimental data of 200 points under varying conditions supported machine learning simulations. Bayesian optimization identified the Gaussian process as the optimal model. The proposed CIML framework integrates the GP prediction (R2 = 0.9979 for Ce and 0.9993 for Qe) with mass balance formulation, error propagation, residual correction stage and uncertainty quantification. To further access generalizability, external validation is performed using 63 independent data points under varying conditions. CIML performance is benchmarked against, conventional mass balance model, predicting Qe directly from Ce using mass balance equation and a two stage GP model lacking physical constraints. The results show that the proposed CIML framework achieves R2 = 0.9993 and RMSE = 0.0493, outperforming both the mass balance model (R2 = 0.9974 and RMSE = 0.0973) and the two stage GP model (R2 = 0.9973, RMSE = 0.0938) corresponding to an approximate 47% reduction in Qe prediction error. Multi-layer validation confirmed robustness, while uncertainty quantification showed 96.8% coverage. Interpretability analysis identified the key influential variables. The results demonstrates that the CIML enhances accuracy, generalizability, and reliability for wastewater treatment applications.

Meghavi.D Parmar, Vipin Shukla, M. E. Ali Mohsin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.