A Multi-UAV Cooperative Mission Planning Method Based on Multi-Agent Guided Soft Actor–Critic
Abstract
Multiple unmanned aerial vehicles (UAVs) performing cooperative missions in complex environments face challenges such as difficult cooperative decision-making, stringent spatiotemporal consistency constraints, and environmental uncertainty. The cooperative mission considered in this paper aims to enable multiple UAVs to simultaneously arrive at multiple constant-velocity moving targets. To address these challenges, this paper proposes a multi-agent guided soft actor–critic (MAGSAC) deep reinforcement learning algorithm. Under the centralized training with decentralized execution (CTDE) framework, a Guider network is introduced to guide the local actor network in learning coordinated strategies, thereby alleviating the non-stationarity of multi-agent decision-making under uncertain environments. An estimated time of arrival (ETA)-based spatiotemporal coordination reward function is designed to promote synchronized arrival. To address sparse rewards, a hindsight experience replay (HER) mechanism based on backward trajectory reconstruction is developed, and a delayed collision-constraint activation mechanism is incorporated to improve convergence while maintaining flight safety. Simulation results show that MAGSAC outperforms existing mainstream algorithms in synchronization success rate, temporal synchronization accuracy, and safety.