Skip to content
Review

HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows

Jul 2026 · arXiv.org · Vol abs/2607.23983 · 0 citations · 65 references
Computer Science Physics

TL;DR

HydroAgent is proposed, a skill-orchestrated agent framework that embeds Large Language Models (LLMs) into a model-driven flood forecasting workflow, where each skill encodes explicit rules to bound LLM reasoning, demonstrating how explicit rule boundaries can guide language model reasoning to complement physically based simulation in next-generation flood forecasting.

Abstract

Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer. Although artificial intelligence methods have advanced flood prediction and model-error correction, most existing studies have not explicitly represented the tacit expert rules, review checkpoints, and workflow constraints that connect model outputs to operational warning decisions. To address this issue, we propose HydroAgent, a skill-orchestrated agent framework that embeds Large Language Models (LLMs) into a model-driven flood forecasting workflow, where each skill encodes explicit rules to bound LLM reasoning. We validated its effectiveness using five state-of-the-art LLMs in the South Yamhill River basin. Our results demonstrate that prior judgment captures observed peak flow and flood volume within 5% tolerance in 10 and 11 out of 14 events, with 5-fold cross-validation over 129 events yielding Pearson correlations of 0.62 and 0.84. Building on a high-baseline scheme library (average KGE 0.890), the guided scheme selection further improves KGE by 0.023-0.154, with simulated peak flow and flood volume falling within the prior judgment ranges for 14 and 13 out of 14 events. All five tested LLMs successfully execute the HydroAgent workflow with comparable judgment accuracy (40%-80%), while showing moderate performance variation and substantial cost differences. HydroAgent does not aim to replace human forecasters; instead, it translates their tacit expertise into an auditable and reproducible workflow, streamlining analytical steps and supporting more informed decision-making. This skill-orchestrated paradigm demonstrates how explicit rule boundaries can guide language model reasoning to complement physically based simulation in next-generation flood forecasting.

View source

Similar papers

Jul 2026

Context-Aware Concept Distillation for Trustworthy Flood Prediction

Context-Aware Concept Distillation (CACD) is proposed, a framework developed in collaboration with domain experts to distill opaque LSTMs into interpretable, hydrology-aware surrogate models and a Residual Hypernetwork that dynamically modulates these concepts based on static basin characteristics.

Eli Levinkopf, E. Morin, Claudia V. Goldman · 0 citations
Open access Sep 2026

Uncertainty-aware maintenance scheduling in water distribution networks via ensemble neural forecasting and explainable confidence indexing

Maintenance scheduling in water distribution networks (WDNs) needs accurate demand forecasts. However, climate variability makes point predictions insufficient. Existing Digital Twin (DT) frameworks use deterministic forecasts, leading to over-scheduling and Service Level Agreement (SLA) violations during volatile weather. Bayesian uncertainty methods are rigorous but require 340 ms per inference, making them too slow for real-time scheduling on standard utility hardware. We propose CAUCCES, coupling an adaptive ensemble (LSTM, Prophet, LightGBM, XGBoost) with a novel Explainable Confidence Index (ECI). ECI is a closed-form uncertainty measure based on ensemble entropy and variance. It directly connects the forecasting module to a constraint-based scheduler. When confidence drops, non-critical tasks are deferred, turning forecast uncertainty into actionable decisions. Validated across 12 Spanish municipalities over 18 months, ECI-driven scheduling reduces SLA violations from 9.3% to 1.5% with only 3.1% extra operational cost. The ensemble achieves 14.12% MAPE, outperforming DeepAR (16.50%) and Temporal Fusion Transformer (16.92%) Furthermore, it achieves a 28\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\times$$\end{document} speedup (12 ms vs. 340 ms) compared to a Bayesian LSTM baseline. Finally, multi-regional testing shows 12–26% performance degradation across different climatic zones, showing that local recalibration is necessary. These results suggest that entropy-based uncertainty quantification can serve as a practical bridge between forecasting and scheduling for resource-constrained utilities, although broader validation across climates and operating conditions is still needed.

MohammadHossein Homaei, Óscar Mogollón-Gutiérrez, Mostafa M. Rezaee et al. · 0 citations
Open access Aug 2026

Automating SWMM-based stormwater modelling and analysis through a tool-augmented single-agent system.

Urban stormwater modelling plays a critical role in assessing interventions for flood risk and water quality management in response to ageing infrastructure and future uncertainties. However, modelling workflows in practice remain highly manual, and key steps in model configuration, execution, and interpretation often depend on specialised knowledge, leading to inefficiencies. Therefore, this study proposes SWMM-Agentic, a tool-augmented, large language model (LLM)-based single-agent system for urban stormwater modelling, simulation, and scenario analysis. Built on the Storm Water Management Model (SWMM), SWMM-Agentic uses one orchestration model to interpret natural-language instructions and sequentially invoke documented functions for traceable post-configuration workflows. Evaluation on the Astlingen benchmark included capability demonstrations and a 60-task suite comprising 20 static, 20 dynamic, and 20 scenario-based tasks, executed once with each of three LLMs to produce 180 model-task runs. DeepSeek-V3.2-Exp successfully completed 59/60 tasks (98.3%), Qwen3-236B completed 58/60 (96.7%), and Qwen3-14B completed 45/60 (75.0%). Across 180 runs, 89 of 100 failed tool calls were followed by a successful corrective call within three attempts. SWMM-Agentic also reproduced network characteristics, compared alternative control strategies, and conducted a human-framed rain-garden experiment that showed decreasing combined sewer overflow discharge with diminishing marginal benefits at higher coverage. These results demonstrate that SWMM-Agentic can reliably operate existing SWMM models through natural language within the evaluated benchmark and tool scope, supporting accurate and reproducible stormwater simulation and analysis, and laying the groundwork for natural-language-driven platforms for integrated planning and hypothesis-driven research.

Jian Wang, Chenyue Sun, Dragan A. Savić et al. · 0 citations
Review Open access Aug 2026

Artificial Intelligence for Hydraulic-Fracturing Decision Support: A Workflow-Oriented Critical Review

Hydraulic fracturing is a critical technology for unconventional oil and gas development, but its performance is strongly affected by geological heterogeneity, complex fracture propagation, operational uncertainty, and nonlinear interactions among engineering parameters. Artificial intelligence (AI) provides tools for extracting relationships from geological, geophysical, operational, and production data. This structured narrative review synthesizes AI applications across four sequential stages of the hydraulic-fracturing workflow: sweet-spot identification, fracturing-parameter optimization, operational diagnosis and risk warning, and post-fracturing flowback prediction and control. Representative studies reported sweet-spot classification accuracy of 97.5% and R2 = 0.97 for production-performance prediction; a simulator-coupled optimization study reported a 13% economic improvement, and field-data models used cohorts of up to 295 wells. Operational studies reported point-event recognition above 97%, pressure forecasting 30 s ahead, and risk forecasts over three consecutive 60 s intervals. A post-fracturing model trained on 286 wells predicted responses over 30-, 90-, 180-, and 360-day horizons. These values are study-specific and are not directly comparable because the datasets, targets, partitions, and metrics differ. Collectively, the evidence indicates measurable but uneven progress; field readiness remains limited by data quality, multimodal alignment, physical consistency, uncertainty quantification, external validation, and weak coupling between model outputs and operational decisions. The review contributes a reproducible workflow-oriented coding framework and defines validation and deployment priorities for reliable, interpretable, and executable AI-assisted fracturing decision support.

Xiao-Bing Bian, Jia-Xing Zhou, Liang Fu et al. · 0 citations
Preprint Aug 2026

REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting

REATS is proposed, which leverages LLM reasoning capabilities as an intelligent ensemble router that jointly processes textual temporal pattern descriptions and numerical features to produce interpretable, sample-adaptive ensemble weights through chain-of-thought reasoning.

Xu Zhang, Chang Xu, Hui Sun et al. · 0 citations
Open access Sep 2026

Riverine Flood Forecasting Using Advanced Deep Learning Approaches

Accurate riverine flood forecasting is crucial for effective river management. This paper utilized the Time‐Series Dense Encoder (TiDE), Neural Hierarchical Interpolation for Time Series Forecasting (N‐HiTS), and Patch Time Series Transformer (PatchTST) to forecast riverine flood and benchmarked their results against Long Short‐Term Memory (LSTM). Each model was implemented with varying forecast lead times for the Proctor Creek–Chattahoochee watershed, Georgia, USA. Additionally, the sensitivity of each model was evaluated by excluding meteorological forcing features one by one to determine how the performance varied across different variables. The time of concentration () was incorporated as a physical parameter in the algorithm's lookback window. The trained models were evaluated separately on event‐based simulations. The Diebold–Mariano statistical test was utilized for a thorough analysis of performance. Analysis revealed that PatchTST outperformed the other models during moderate‐flow regimes while falling behind during extreme flooding events. An 18‐ to 24‐h lookback window was found to be statistically optimal for the models. The sensitivity analysis results indicated that PatchTST was slightly more sensitive to the selected training data features. Incorporating into the lookback window revealed that TiDE and LSTM showed better performance when using a sequence length of 18 h while PatchTST achieved the best results with a 24‐h lookback window.

Krishna Panthi, Mostafa Saberian, Vidya S. Samadi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.