Skip to content
Open access

Modeling Dynamic Obstacle Avoidance Strategy of Drone Swarms Combined with Multi-Agent Reinforcement Learning

Aug 2026 · Advanced Electromagnetics · 0 citations

TL;DR

The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.

Abstract

This paper proposes the Locally-decoupled and Embedding-enhanced Multi-Agent Deep Deterministic Policy Gradient (LDE-MADDPG) algorithm to address poor scalability and delayed response in drone swarm dynamic obstacle avoidance under complex cooperative environments. Such autonomous coordination capabilities are also important for distributed sensing, wireless networking, and electromagnetic information exchange in future intelligent aerial systems. The algorithm introduces three key innovations beyond standard MADDPG: a Graph Attention Network module that encodes variable-length observations into fixed-dimensional embeddings for swarm-size generalization; a dual-path critic with a global branch guiding policy updates and a local branch specializing in obstacle avoidance evaluation; and a hierarchical reward integrating multi-objective signals. Evaluated across eight static and dynamic obstacle scenarios, LDE-MADDPG achieves significantly lower collision rates (2.1%–4.2% in static scenarios and 3.8%–7.2% in dynamic scenarios) than state-of-the-art baselines and reaches a 97.5% mission completion rate in 100 random scenarios. The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.

Read PDF

Similar papers

Conference Aug 2026

Attention-based MADDPG with dual-buffer experience replay for cooperative multi-UAV target tracking

An Attention-based Multi-Agent Deep Deterministic Policy Gradient algorithm was developed for cooperative multi-unmanned aerial vehicle target tracking in dynamic environments. The study addressed information redundancy and association weight allocation between individual unmanned aerial vehicles and the swarm during cooperative decision-making. To improve information selection, the proposed algorithm introduced a centralized critic network with a multi-head attention mechanism to evaluate the contributions of other agents at each time step. Meanwhile, the study designed a dual-buffer experience management architecture composed of a recent interaction memory and a mission outcome memory. This architecture stored recent interaction data and mission-critical trajectories separately, thereby improving experience utilization during training. The study also formulated the state space, action space, and reward function for target exploration, obstacle avoidance, energy consumption, and velocity maintenance under limited perception conditions. The experiments evaluated the proposed algorithm in a two-dimensional tracking scenario with moving targets, multiple unmanned aerial vehicles, and dynamic and static obstacles. The proposed method was compared with Deep Deterministic Policy Gradient and Multi-Agent Deep Deterministic Policy Gradient using collision rate, capture rate, capture time, and capture distance as evaluation metrics. The results showed that the proposed algorithm improved convergence behavior and average reward, reduced the collision rate by 83% compared with Deep Deterministic Policy Gradient at the first environmental level, and maintained competitive performance in capture rate, capture time, and path efficiency.

Qinglin Han, Hongmei Wang · 0 citations
Open access Aug 2026

Directional Pheromone Gradient Observations for Decentralized Multi-Agent Reinforcement Learning in Swarm Drone Search and Rescue

Findings indicate that directional pheromone-gradient observations provide an effective and communication-efficient mechanism for decentralized swarm coordination, improving search effectiveness and operational robustness in post-disaster SAR scenarios.

Peter Yacoub, Mohamed Malek Kaouach, Esraa Khatab et al. · 0 citations
Open access Aug 2026

Spatio-Temporal Attention-Based Improved MADDPG Algorithm for Multi-UAV Formation Path Planning

With the increasing deployment of multi-unmanned aerial vehicle (multi-UAV) systems in dynamic environments, the problem of efficient cooperative path planning has emerged as a critical challenge requiring urgent solutions. To address this issue, this paper proposes a novel joint optimization framework, named spatio-temporal attention-based multi-agent deep deterministic policy gradient (STA-MADDPG). Rather than proposing a new reinforcement learning algorithm in the strict sense, this work integrates advanced spatial-temporal feature extraction with heuristic gradient guidance. First, a cascaded architecture combining multi-head attention and Long Short-Term Memory (LSTM) networks is utilized to extract key local and temporal features, thereby mitigating the dimensionality curse in dense multi-agent observations. Second, an improved dynamic artificial potential field (DAPF) is integrated into the reinforcement learning framework as a state augmentation mechanism, providing heuristic guidance vectors that accelerate convergence and improve obstacle avoidance. Furthermore, to balance computational complexity and adaptive behavior, a rule-based hierarchical formation strategy is designed. The framework maps predefined formations (elliptical, chain, or wedge) to specific environment categories, while the underlying MARL policy governs the dynamic trajectory planning and topology maintenance. Finally, rigorous comparative and ablation experiments are conducted to evaluate path length, search time, and relative position errors. Statistical analysis demonstrates the effectiveness of the proposed framework, achieving up to a 67.3% reduction in search time and a 91.56% search success rate compared with standard MARL baselines in complex environments.

Dong Zhao, Huaizhi Dong, Wenjing Ren · 0 citations
Review Jul 2026

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

A multi-agent deep reinforcement learning framework that addresses issues through coordinated exploration, demonstration exploitation, safe curriculum scheduling, and structure-aware generalisation is proposed, demonstrating strong performance in collaboration success rate, navigation robustness, zero-shot cross-scenario generalisation, and dynamic environment adaptability.

Yuhuang Su, Nabil Aouf · 0 citations
Aug 2026

Deep reinforcement learning–based safe path planning for leader–follower robots

This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.

Ehsan Kazemi Tameh, Mohammadreza Estarki, Saeed Khodaygan · 0 citations
Conference Sep 2026

Intelligent autonomous navigation for UAVs: a strategic hierarchical path planning framework

To address the challenge that single algorithms struggle to balance global exploration and local obstacle avoidance, and are prone to falling into local optima in complex environments, this paper proposes a Strategic Hierarchical Path Planning (SHPP) framework. This framework decouples the 3D navigation task into three synergistic layers. The top layer employs a reinforcement learning network equipped with a Credit Alignment Mechanism (CAM) to provide macroscopic guidance, eliminating credit assignment pollution and escaping local minima. The middle layer introduces a Dual-Guided Particle Swarm Optimization (DG-PSO) algorithm to map discrete commands into continuous smooth trajectories. The bottom layer executes physical collision avoidance and tracking based on the Artificial Potential Field (APF) method. Simulations indicate that the system can establish a stable policy in approximately 625 episodes, achieving an average reward of 95.6. Furthermore, benefiting from the hierarchical architecture's smooth optimization in continuous space, the average flight path length is 138.3 meters, a reduction of approximately 19% compared to traditional discrete decision-making models. These quantitative results fully validate the superior performance of the proposed architecture in complex 3D environments.

Fei Wang, Jun-Yong Shi, Zhao-Kun Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.