Skip to content
Open access

Refined simulation of groundwater level dynamics based on a machine learning stacking framework

Aug 2026 · Frontiers in Water · Vol 8 · 0 citations · 55 references

Abstract

Accurate and rapid prediction of groundwater levels (GWL) is essential for effective groundwater management. Machine learning models are efficient tools for GWL prediction, but individual models often suffer from limited generalization due to inherent randomness. This study proposed a stacking-based GWL prediction framework suitable for arid regions in Northwest China. Feature variables affecting GWL were selected using variable importance in projection (VIP). Then, three machine learning models—artificial neural networks (ANN), random forests (RF), and Light Gradient Boosting Machine (LightGBM)—were developed, and their outputs were integrated using a support vector regression (SVR)-based stacking method to enhance the accuracy of GWL prediction. The results show that the factors of influencing GWL changes vary significantly across different regions, and selecting the most contributive feature variables is beneficial for model construction. Among the individual models, the RF model demonstrated higher accuracy and more stable performance, outperforming the ANN and LightGBM models. However, individual models exhibited poor generalization during validation. In contrast, the stacking model maintained high performance, demonstrating superior generalization. Compared to the best-performing individual model (RF) in validation period, the Nash–Sutcliffe efficiency ( NSE ) and Kling–Gupta efficiency ( KGE ) of stacking model improved by 0.11–0.66 and 0.05–0.41, the correlation coefficient ( R 2 ) increased by 0.05–0.3, and root mean square error ( RMSE ) reduced by 0.01–0.1 m. In the stacking simulation, RF had the highest average contribution (80.2%), followed by ANN (13.9%) and LightGBM (5.9%). This study provides a stacking simulation framework based on machine learning methods for precise groundwater level simulation, which can serve as a reference for groundwater level simulation in other regions.

Read PDF

Similar papers

Open access Sep 2026

Comparative assessment of machine learning models for seasonal groundwater-level prediction in the Hiran Basin, India

Groundwater monitoring and water-resource management in the context of growing climatic variability requires accurate prediction of groundwater levels. In this study, six machine-learning models were tested including Random Forest (RF), Extreme Gradient Boosting (XGB), Extra Trees Regressor (ETR), Histogram Gradient Boosting (HGB), Artificial Neural Network (ANN), and Support Vector Regression (SVR) models to predict seasonal groundwater levels in the Hiran Basin using hydroclimatic and antecedent groundwater-level variables. Precipitation, maximum temperature, minimum temperature, and four groundwater-level variables that were lagged by varying amounts were used in developing the models, and these variables were obtained from long-term observations of groundwater levels in monitoring wells throughout the basin. The performance of the models was evaluated by using a chronological validation framework and comparing the predictive ability of the six algorithms. The analysis revealed that the ANN model performed best in predicting overall results with HGB, XGB, RF, SVR, and ETR also giving reliable predictions. Precipitation and antecedent groundwater levels were the most important variables for explaining groundwater-level change in the SHAP analysis. The proposed framework is practical for seasonal groundwater-level forecasting and can be applicable for monitoring groundwater levels and managing groundwater resources in the Hiran Basin.

Unknown authors · 0 citations
Open access Aug 2026

Advancing River Water Level Prediction: A Comparative Machine Learning and Deep Learning Approach

River water level prediction plays an important role in effective planning and flood risk mitigation. In this study, four standalone machine learning (ML) models, M5Rules, Random Forest (RF), Sequential Minimal Optimization (SMO), and Long Short-Term Memory (LSTM), as well as a hybrid LSTM-RF model, were developed to predict weekly water levels of the Rhine River. The models were trained and tested using data collected between 2006 and 2024. Different scenarios with different input combinations were explored to improve the accuracy of the model. Statistical indicators were calculated to examine the reliability of the proposed scenarios and models. The results showed that the performance of the model increased in Scenario 4 with all input variables. Among the standalone models the M5Rule and SMO algorithms perform better with Nash-Sutcliffe Efficiency (NSE) of 0.78, in validation phase, followed by RF (NSE = 0.76) and LSTM (NSE = 0.72). In order to increase the model predictive power, the hybrid model LSTM-RF applied to the input variables of the best scenario and this hybrid model achieved a remarkable accuracy of NSE = 0.98 significantly outperforming standalone models. The findings of this research demonstrated the efficacy of the hybrid LSTM-RF model in capturing the changes in the water level in Rhine River.

Zohreh Sheikh Khozani, Monica Ionita · 0 citations
Open access Sep 2026

A Hybrid Ensemble Learning Approach for Accurate SPI₆-Based Drought Forecasting in Semi-Arid Regions

This study proposes a hybrid machine learning framework to predict the six-month Standardized Precipitation Index (SPI₆) for meteorological drought assessment in Nanded, India, using NASA POWER data (1994–2024). Four models Ridge Regression, Random Forest, Multi-Layer Perceptron (MLP) and a stacking ensemble (Ensemble_StackLR) were developed and evaluated using R², RMSE, MAE, NSE and PBIAS. Random Forest and MLP showed strong predictive capability, while Ridge Regression provided stable but comparatively lower performance in capturing nonlinear patterns. The Ensemble_StackLR model delivered the most robust and balanced results, achieving R² between 0.87 and 0.91, high correlation (r ≈ 0.96) and minimal bias across training, validation and testing datasets. It effectively captured drought onset, duration and recovery phases, outperforming individual models in stability and generalization. The framework demonstrates that ensemble learning enhances SPI prediction accuracy and offers a scalable, data-driven solution for drought monitoring and water resource management in data-scarce regions.

Unknown authors · 0 citations
Open access Aug 2026

Interpretable application of machine learning techniques in the assessment of coastal water quality of Poyang Lake, China

Machine learning models (MLMs) have made substantial progress across diverse scientific domains owing to their strong predictive capabilities and effectiveness in extracting patterns from high-dimensional data. In lake environmental management, clarifying the relationships between environmental factors and water quality indicators (WQIs) is essential for improving assessment accuracy and operational efficiency. In this study, four MLMs (support vector regression [SVR], random forest [RF], extreme gradient boosting [XGBoost], and k-nearest neighbors [KNN]) were applied to identify the key environmental parameters influencing WQIs in the coastal zone of Poyang Lake, China, and to evaluate the predictive performance of each model. Model interpretability was enhanced using the SHapley Additive exPlanations (SHAP) method, which quantified the contributions of environmental variables to WQI variability. Total nitrogen (TN), total phosphorus (TP), and chlorophyll-a (Chla) concentrations were used as representative WQIs. Among all models, XGBoost consistently exhibited the highest predictive accuracy. Electrical conductivity (EC), water temperature (T), dissolved oxygen (DO), oxidation–reduction potential (ORP), and pH were identified as the five most influential parameters. Specifically, EC, T, and DO were the dominant drivers of TN and Chla, accounting for 92.8% and 84.2% of the total SHAP contributions, respectively. Meanwhile, TP was primarily governed by ORP, pH, and T, with a combined contribution of 70.6%. XGBoost accurately reproduced the observed WQIs using environmental parameters as inputs. These findings demonstrate that interpretable MLMs can effectively reveal the environmental drivers of lake water quality and offer a robust framework for ecosystem monitoring and forecasting in freshwater environments.

Zhang Hu, Rong Yi, Xin Liu et al. · 0 citations
Open access Jul 2026

Integrating SWAT and machine learning for streamflow simulation and runoff sensitivity in the Tawi watershed

Accurate streamflow simulation is essential for water resource management in data-scarce Himalayan watersheds, where hydro-climatic variability and land-use changes influence hydrological processes. This study evaluates the performance of the Soil and Water Assessment Tool (SWAT) and machine learning (ML) models, Random Forest (RF), Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Artificial Neural Networks (ANN), for simulating streamflow in the Tawi River watershed using hydroclimatic data from 2000 to 2020. SWAT performance was assessed using statistical indicators and hydrograph analysis. Despite a negative PBIAS (− 30.92%) during calibration (2002–2014), the model showed satisfactory performance (R2 = 0.80, NSE = 0.67) and improved validation results (2015–2020) (R2 = 0.85, NSE = 0.77, PBIAS = − 12.50%). Machine learning models were trained on historical data and evaluated over the overlap period (2015–2020). XGBoost (R2 = 0.77, RMSE = 24.34 m3/s) and Random Forest (R2 = 0.74) outperformed SVR and ANN, which showed lower accuracy and underestimation of peak flows. Comparative analysis indicates that SWAT provides physically interpretable and consistent simulations of streamflow dynamics, whereas machine learning models, particularly XGBoost and RF, offer efficient data-driven predictions with strong capability in capturing nonlinear relationships. Sensitivity analysis revealed rainfall as the dominant control on streamflow (εRF up to 4.93), while temperature showed a comparatively weaker influence. Overall, the results demonstrate that process-based and data-driven approaches provide complementary strengths for streamflow simulation; however, the extended simulations (2021–2050) represent statistical or stationary projections and should not be interpreted as climate-driven future forecasts.

A. S. Jasrotia, Komal Kumar Singh, Praveen Thakur et al. · 0 citations
Open access Jul 2026

Comparative evaluation of AHP and advanced machine learning models for delineating groundwater potential zones in Southern region of Delhi, India

Groundwater resources in the southern region of Delhi are under severe stress due to rapid urbanization, excessive abstraction, and declining recharge. This study delineated groundwater potential zones (GWPZs) by comparing a conventional Analytical Hierarchy Process (AHP) model with four data-driven approaches: XGBoost, Random Forest (RF), Multilayer Perceptron Neural Network (MLPNN), and Deep Learning Neural Network (DLNN). Thirteen conditioning factors representing geology, geomorphology, hydrology, terrain, and land-surface characteristics were integrated within a common GIS-based framework. Model performance was evaluated using internal hold-out ROC-AUC analysis and threshold-based confusion matrices and was further examined through independent field validation using 51 georeferenced wells. The DLNN model achieved the highest predictive accuracy (AUC = 0.932), followed by MLPNN (0.901), XGBoost (0.891), RF (0.871), and AHP (0.824), while external validation preserved the same overall ranking. High to Very High groundwater potential zones were concentrated mainly along the Yamuna floodplain and favorable alluvial sectors, whereas the central and southwestern hard-rock and densely urbanized areas were dominated by lower potential classes. The study demonstrates that data-driven models, particularly DLNN and MLPNN, provide more reliable delineation of GWPZs than the knowledge-driven AHP baseline and that independent field validation strengthens the practical value of the generated maps for groundwater management and recharge planning in rapidly urbanizing environments.

Deepanshi Tanwar, K. Sarma, Sayantan Mandal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.