Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Dual-Agent Hierarchical Reinforcement Learning for Typhoon-Avoidance Route Planning of Ships Under Dynamic Wind–Wave–Current Fields

Ship weather routing in regional seas under severe seasonal weather systems, such as typhoons, poses a critical operational challenge for maritime safety and efficiency. Traditional single-agent reinforcement learning (RL) methods frequently suffer from training instabilities and erratic trajectory adjustments when exposed to the highly non-stationary, multi-scale dynamics of coupled wind–wave–current fields. To address these limitations, this study proposes a dual-agent hierarchical reinforcement learning framework (HRL-MOO) featuring a built-in dynamic risk-weight adaptation mechanism. The proposed global agent discerns large-scale environmental evolutions and adaptively updates the relative weights of wind-, wave-, and current-induced risks at a strategic level, while the local agent translates this macro-level guidance into short-term tactical heading and speed adjustments within realistic vessel motion boundaries. The framework incorporates bathymetric constraints through a high-resolution navigable domain mask derived from ETOPO topography to guarantee practical navigability. Simulated experiments are executed using hourly environmental reanalysis data and best-track records corresponding to the passage of Typhoon Yagi (2024) through the northern South China Sea and Taiwan Strait. The empirical results demonstrate that the cooperative dual-agent structure establishes a superior global Pareto frontier compared to conventional standard single-agent Proximal Policy Optimization (PPO) and metaheuristic baselines. Under extreme typhoon conditions, the HRL-MOO model effectively decreases the cumulative environment-induced risk from approximately 0.17 to 0.12, improves path smoothness by achieving a higher index of 0.842, and accelerates policy convergence to within 3000 training episodes, all while maintaining a highly efficient voyage distance with minimal detour overhead. Ablation studies further validate that the synergy between macro-level stream field recognition and adaptive multi-objective optimization significantly enhances decision-making stability. This framework offers a robust and interpretable computational tool for autonomous ship weather routing in fast-changing, high-risk ocean environments.

Yu Cai, Ying Li, Liankang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.