Hierarchical Deep Reinforcement Learning with Reward Shaping for Fixed-Wing Multi-UAV Trajectory Tracking and Collision Avoidance
Abstract
Trajectory tracking for fixed-wing unmanned aerials (UAVs) in obstacle-cluttered environments remains a formidable challenge due to their inherent non-holonomic kinematic constraints. This problem is further exacerbated in dense multiUAV scenarios, where the system must navigate complex spatial geometries to simultaneously execute static obstacle avoidance and dynamic inter-agent collision avoidance. To address these tightly coupled objectives, this paper proposes a novel hierarchical hybrid control framework that seamlessly integrates deep reinforcement learning with the L1 navigation algorithm. A key contribution of this work is the development of a continuous adaptive reward shaping mechanism. Within this hierarchical architecture, the soft actor-critic agent functions as a high-level planner generating virtual path offsets, while the L1 controller ensures smooth, kinematically feasible low-level execution. Extensive simulations demonstrate that the proposed method ensures high-fidelity trajectory tracking, allowing UAVs to execute smooth evasive maneuvers and rapidly return to the reference path, thus guaranteeing safety and efficiency in highly demanding multi-agent intersections and curved scenarios.