Hierarchical Multi-Agent Reinforcement Learning for Networked Multi-AUV Data Collection in UWSNs
Abstract
Underwater wireless sensor networks form the foundation of the marine Internet of Things, but timely data delivery remains challenging in deep and remote deployments. Although AUV-assisted data collection reduces reliance on energy-constrained multi-hop acoustic relays, slow vehicle mobility and repeated surfacing for data upload still degrade the Value of Information (VoI). To address this bottleneck, we propose a networked multi-AUV data-collection framework supporting opportunistic inter-AUV acoustic relaying. Buffered data can be forwarded to near-surface peers acting as mobile gateways, reducing redundant surfacing and improving delivered VoI at the surface base station. We formulate a VoI-driven cooperative control problem jointly considering sensor-specific acquisition, peer-specific relaying, uploading, idling, and continuous three-dimensional trajectory control under communication-feasibility and collision-avoidance constraints. To learn a tractable policy for this strongly coupled hybrid decision problem, we develop a hierarchical multi-agent reinforcement learning framework. The upper layer uses QMIX-style value decomposition to learn coordinated discrete task decisions from local observation histories, while the lower layer uses MADDPG with centralized critics and decentralized actors to learn task-conditioned continuous three-dimensional motion policies. The two layers are coupled through discrete-action embedding, a shared team reward, and the post-resolution environment transition, avoiding direct optimization over the full joint hybrid action space. Extensive simulations demonstrate competitive delivered-VoI performance and faster, more stable convergence than representative baselines, while substantially reducing surfacing frequency of deep-water collectors.