Attention-Enhanced Hierarchical Reinforcement Learning for Air-Ground Cooperative Perception
Abstract
Air-ground cooperative perception (AGCP) integrates connected and autonomous vehicles (CAVs), roadside units (RSUs), and uncrewed aerial vehicles (UAVs) to provide wide-area coverage and high-resolution perception by leveraging their complementary perception and communication capabilities. However, the dynamic and heterogeneous characteristics of the air-ground network introduce strong cross-layer coupling across perception, communication, and computation, thereby complicating the coordination of cooperation update intervals and cooperation partner selection. To address these challenges, we develop a unified AGCP framework that jointly models LiDAR-based multi-agent perception, together with its associated communication bandwidth allocation and computation latency models, under dynamic mobility and time-varying resource conditions. Building on this framework, a multi-objective optimization problem is formulated to characterize the interplay between update interval selection and cooperation partner choice, aiming to balance perception accuracy and end-to-end latency. A Tchebycheff distance-based formulation is utilized to normalize and integrate multiple objectives into a unified optimization metric. To efficiently solve this highly coupled problem, an attention-enhanced hierarchical reinforcement learning algorithm is proposed, which leverages a two-level Markov decision process combined with an attention-enhanced actor-critic architecture. Simulation results validate that the proposed algorithm achieves a desirable trade-off between perception performance and end-to-end latency.