Similar papers
Understanding Long-Term Groundwater Storage Variability Using GRACE Data and Explainable Machine Learning
The decline in groundwater storage (GWS) poses a critical threat to water security in semi-arid regions where increasing agricultural water demand and climate variability are increasing pressure on aquifers. This study presents a novel hybrid modeling framework integrating multi-source satellite and climate data (GRACE, GLDAS, TerraClimate, and MODIS) with machine learning and explanatory artificial intelligence techniques for the long-term assessment and interpretation of GWS anomalies in the data-poor Iğdır Basin. Three different modeling approaches were developed: XGBoost, Long Short-Term Memory (LSTM) networks, and their combined model, and interpreted using the Shapley Additive Explanations (SHAP) method. The results showed a significant long-term decreasing trend in groundwater storage anomalies at a rate of −0.87 mm per month during the 2002–2016 period, indicating continuous depletion. The LSTM model demonstrated the best performance with R2 of 0.59, RMSE of 19.5 mm, and MAE of 15.1 mm, revealing the dominant role of temporal dependencies in groundwater systems. SHAP analysis identified lagged groundwater anomalies (especially GWS_lag3) as the most effective predictors; this may reflect the memory effect and lagged response specific to semi-arid aquifer systems, but this interpretation needs to be validated in different study areas. Snow water equivalent and total water storage anomalies also emerged as significant determinants, while the direct effect of instantaneous precipitation was found to be limited. This study addresses significant gaps in the literature by combining sequence-based modeling with model interpretability in a semi-arid closed basin. The findings highlight the necessity of using system memory and explainable artificial intelligence together for reliable groundwater prediction. While the proposed hybrid approach has the potential for application in other semi-arid regions, its broader usability needs to be supported by independent validation studies under different hydrogeological and climatic conditions.
Machine Learning-Based Assessment of Long-Term Groundwater Level Dynamics in the Nakhon Luang Aquifer from 1978 to 2023
Groundwater is a critical resource supporting economic and social development in Thailand, particularly within the lower Chao Phraya Basin, where rapid urbanization and industrial expansion have led to intensive groundwater exploitation. This study applied a random forest (RF) machine learning model to investigate groundwater level (GWL) dynamics in the Nakhon Luang aquifer, which is one of the most productive aquifers due to its high yield and relatively good water quality, using long-term data from 1978 to 2023. The seven data sets include groundwater levels, groundwater pumping, rainfall, groundwater recharge, ground surface elevation, geological information, and lag-time features incorporated to represent delayed aquifer responses. The results demonstrate strong predictive performance, with R2 = 0.851 for the testing data set. Cross-validation results further indicated stable model performance (R2 = 0.833 ± 0.018). Feature importance analysis revealed that groundwater pumping and recharge lag variables were the most influential factors controlling groundwater level variations. In addition, residual analysis and spatial error mapping were conducted to evaluate model uncertainty, while SHAP analysis was used to interpret the influence of input variables on groundwater level predictions. This study provides new insights into the dynamics of confined aquifer systems using machine-learning techniques and offers valuable information to support sustainable groundwater management in the study area.
An End-to-End Machine Learning Framework for Groundwater Level Characterization and Climate-Constrained Probabilistic Forecasting in a Complex Karst Aquifer
Human activities such as intensive groundwater abstraction and mine dewatering can profoundly disrupt the natural hydrological functioning of karst aquifers. The Transdanubian karst aquifer in Hungary represents one of Central Europe’s most prominent examples, where decades of coal-mine dewatering lowered groundwater levels by more than 40 m and fundamentally altered the natural recharge–discharge regime. Understanding and forecasting recovery in such complex karst systems remain challenging because of heterogeneous conduit–fracture networks, strong climate sensitivity, incomplete monitoring records, and uncertainty in long-term predictions. This study presents an integrated end-to-end machine learning framework for groundwater characterization and climate-constrained probabilistic forecasting. Monthly groundwater-level records (1970–2026) from five monitoring wells were first reconstructed using a hybrid Moving Average–Random Forest gap-filling approach, achieving high reconstruction accuracy (R2 = 0.87–0.98). Self-Organizing Maps subsequently identified four hydrogeological states representing the dewatering, transition, recovery, and near-equilibrium phases, while inter-well weight-plane correlations (>0.95) confirmed strong basin-scale hydraulic connectivity. A Bootstrapped Random Forest model forced by bias-corrected COSMO-CLM precipitation projections under the SSP2-4.5 climate scenario generated probabilistic groundwater forecasts through 2030, achieving high predictive performance (NSE > 0.80; RMSE = 0.10–0.35 m). Forecast results indicate that the basin as a whole is approaching hydraulic equilibrium by 2030, with distinct well-specific trajectories including mild steady decline and near-stable water level. The proposed framework provides a robust and transferable methodology for groundwater characterization and long-term forecasting in complex karst and fractured aquifer systems under changing climatic conditions.
Advancing sustainable groundwater mapping and management in arid quaternary aquifers using machine learning and geospatial analytics integrating remote sensing and field hydrogeological data
Groundwater (GW) represents a critical resource for sustaining agriculture and rural communities across the arid regions of many developing countries. This study assesses three predictive approaches boosted classification tree (BCT), the bivariate frequency ratio (FR), and a hybrid BCT–FR ensemble for mapping Potential Zones (GWPZ) in arid environments. The modelling framework integrates satellite-derived variables with pumping-test measurements (specific capacity (SPC) and transmissivity (T)) and incorporates topographic, geological, hydrogeological, and anthropogenic factors using an inventory of forty-two wells divided into calibration (70%) and validation (30%) datasets across the West El-Minia region of Upper Egypt. Change-detection analysis over the study period (2000–2025) indicated a substantial increase in agricultural activity, with cultivated lands expanding by more than 560 km². This expansion was accompanied by an observed level decline of approximately five meters over the same period, based on field measurements from 42 wells and calculated using observed water table differences. Based on SPC predictions, the BCT and hybrid FR–BCT models achieved relatively high area under the curve (AUC) values. For transmissivity, the corresponding accuracy values were 83.38% for BCT and 92.58% for FR–BCT. Model outputs were assessing their reliability by comparing the GW potential map generated with available borehole information and daily GWproductivity data from the aquifer system. The study area was classified into four GW potential categories: very high (8%), high (26%), moderate (54%), and low (12%), with the northeastern sector exhibiting the highest recharge and storage potential. Overall, the applied machine-learning techniques demonstrated good performance for GW potential assessment in data-limited environments. The results provide important guidance for GW resource management by identifying zones with substantial recharge and development potential.
Hydrochemical Evolution and Predictive Modeling of Chloride-Dominated Groundwater Salinization in Northern Kuwait.
Groundwater quality assessment in hyper arid aquifers is often limited by sparse monitoring, incomplete hydraulic data, and difficulty separating prior hydrochemical conditions from short-term meteorological forcing. This study develops a hydrochemically grounded and machine learning framework with internal transferability testing to diagnose chloride-dominated salinization in the Rawdatain-Umm Al Aish freshwater aquifers of northern Kuwait. The framework combines four components: comparison of pre-1990 baseline chemistry with 2012-2015 observations, chloride-specific salinization diagnostics, separation of predictors into meteorological, space-time, prior-chemistry, and hybrid groups, and model evaluation using random, temporal, and grouped-well validation. The analysis also integrates spatial salinity mapping, Cl-TDS analysis, Na:Cl ratios, carbonate/sulfate balance, principal component analysis (PCA), and supervised learning models to evaluate both hydrochemical change and predictive structure. This integrated analysis shows substantial deterioration relative to the historical baseline. Median TDS increased from 905.5 to 1821.0 mg/L (+101.1%), whereas median Cl, Ca, Mg, and Na increased by 228.5%, 194.6%, 227.0%, and 75.3%, respectively. HCO3 remained broadly stable, whereas NO3 declined by 43.5%. Diagnostic plots and PCA indicate chloride-rich salinization governed by mixed geochemical and mixing processes rather than simple halite dissolution alone. Random holdout models performed well, with R2 values of 0.947 for SO4, 0.913 for Na, 0.912 for Ca, 0.877 for HCO3, 0.868 for TDS, 0.808 for NO3, 0.807 for Cl, and 0.673 for Mg. However, validation performance was analyte-specific: Cl declined mainly under temporal holdout, NO3 declined under grouped-well holdout, and Mg showed the lowest random-holdout skill and largest train-test gap despite higher grouped-well performance. Predictor importance showed that spatial coordinates, elevation, and lagged chemistry were more influential than meteorological predictors alone. The framework supports targeted monitoring of chloride-rich hotspots and validation-aware groundwater quality decision support in hyper arid freshwater reserves.
Integrating SWAT and machine learning for streamflow simulation and runoff sensitivity in the Tawi watershed
Accurate streamflow simulation is essential for water resource management in data-scarce Himalayan watersheds, where hydro-climatic variability and land-use changes influence hydrological processes. This study evaluates the performance of the Soil and Water Assessment Tool (SWAT) and machine learning (ML) models, Random Forest (RF), Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Artificial Neural Networks (ANN), for simulating streamflow in the Tawi River watershed using hydroclimatic data from 2000 to 2020. SWAT performance was assessed using statistical indicators and hydrograph analysis. Despite a negative PBIAS (− 30.92%) during calibration (2002–2014), the model showed satisfactory performance (R2 = 0.80, NSE = 0.67) and improved validation results (2015–2020) (R2 = 0.85, NSE = 0.77, PBIAS = − 12.50%). Machine learning models were trained on historical data and evaluated over the overlap period (2015–2020). XGBoost (R2 = 0.77, RMSE = 24.34 m3/s) and Random Forest (R2 = 0.74) outperformed SVR and ANN, which showed lower accuracy and underestimation of peak flows. Comparative analysis indicates that SWAT provides physically interpretable and consistent simulations of streamflow dynamics, whereas machine learning models, particularly XGBoost and RF, offer efficient data-driven predictions with strong capability in capturing nonlinear relationships. Sensitivity analysis revealed rainfall as the dominant control on streamflow (εRF up to 4.93), while temperature showed a comparatively weaker influence. Overall, the results demonstrate that process-based and data-driven approaches provide complementary strengths for streamflow simulation; however, the extended simulations (2021–2050) represent statistical or stationary projections and should not be interpreted as climate-driven future forecasts.