Skip to content
Open access

Dual-Agent Hierarchical Reinforcement Learning for Typhoon-Avoidance Route Planning of Ships Under Dynamic Wind–Wave–Current Fields

Aug 2026 · Electronics · 0 citations · 20 references

Abstract

Ship weather routing in regional seas under severe seasonal weather systems, such as typhoons, poses a critical operational challenge for maritime safety and efficiency. Traditional single-agent reinforcement learning (RL) methods frequently suffer from training instabilities and erratic trajectory adjustments when exposed to the highly non-stationary, multi-scale dynamics of coupled wind–wave–current fields. To address these limitations, this study proposes a dual-agent hierarchical reinforcement learning framework (HRL-MOO) featuring a built-in dynamic risk-weight adaptation mechanism. The proposed global agent discerns large-scale environmental evolutions and adaptively updates the relative weights of wind-, wave-, and current-induced risks at a strategic level, while the local agent translates this macro-level guidance into short-term tactical heading and speed adjustments within realistic vessel motion boundaries. The framework incorporates bathymetric constraints through a high-resolution navigable domain mask derived from ETOPO topography to guarantee practical navigability. Simulated experiments are executed using hourly environmental reanalysis data and best-track records corresponding to the passage of Typhoon Yagi (2024) through the northern South China Sea and Taiwan Strait. The empirical results demonstrate that the cooperative dual-agent structure establishes a superior global Pareto frontier compared to conventional standard single-agent Proximal Policy Optimization (PPO) and metaheuristic baselines. Under extreme typhoon conditions, the HRL-MOO model effectively decreases the cumulative environment-induced risk from approximately 0.17 to 0.12, improves path smoothness by achieving a higher index of 0.842, and accelerates policy convergence to within 3000 training episodes, all while maintaining a highly efficient voyage distance with minimal detour overhead. Ablation studies further validate that the synergy between macro-level stream field recognition and adaptive multi-objective optimization significantly enhances decision-making stability. This framework offers a robust and interpretable computational tool for autonomous ship weather routing in fast-changing, high-risk ocean environments.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.