Aug 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 33 references
Medicine
TL;DR
Experimental results demonstrate that the Improved Experience Replay Multi-Agent Deep Deterministic Policy Gradient algorithm outperforms other comparison algorithms in terms of convergence speed, training stability, and path-planning performance.
Abstract
Reinforcement learning techniques have been widely applied to multi-UAV cooperative path-planning tasks. However, existing multi-agent reinforcement learning methods are still affected by environmental non-stationarity, cooperation difficulties among agents, and low utilization efficiency of experience samples in complex obstacle environments. These issues often lead to slow convergence and unstable training performance. To address these problems, an Improved Experience Replay Multi-Agent Deep Deterministic Policy Gradient (IER-MADDPG) algorithm is proposed for multi-UAV cooperative path planning. First, a cooperative path-planning model is established under the Centralized Training Distributed Execution framework. Second, a dual-layer replay buffer structure consisting of a global replay buffer and a local replay buffer is designed to preserve both global cooperative information and individual experience. Third, a fusion experience sampling mechanism is introduced by combining prioritized experience replay and random uniform sampling to improve sample utilization efficiency and training stability. Finally, training experiments were conducted in environments with different obstacle configurations to evaluate the proposed method. Experimental results demonstrate that IER-MADDPG outperforms other comparison algorithms in terms of convergence speed, training stability, and path-planning performance.
An efficient way to resolve the curse of dimensionality, improve obstacle avoidance and cooperative formation control of UAVs was found and shows great prospects of practical application in such domains as military operations, search and rescue missions, transport automation and disaster management.
Qadir Talibov· Problems of Information Tech...· 0 citations
A cooperative guidance law based on the experience-guided multi-agent proximal policy optimization (E-MAPPO) algorithm is proposed for multiple unmanned aerial vehicles (UAVs) to track dynamic points of interest in civilian applications and results indicate that the proposed method generalizes well to different types of maneuvering targets.
Hao Xiong, Minghu Tan, Xiaoyu Liu et al.· Drones· 0 citations
An Attention-based Multi-Agent Deep Deterministic Policy Gradient algorithm was developed for cooperative multi-unmanned aerial vehicle target tracking in dynamic environments. The study addressed information redundancy and association weight allocation between individual unmanned aerial vehicles and the swarm during cooperative decision-making. To improve information selection, the proposed algorithm introduced a centralized critic network with a multi-head attention mechanism to evaluate the contributions of other agents at each time step. Meanwhile, the study designed a dual-buffer experience management architecture composed of a recent interaction memory and a mission outcome memory. This architecture stored recent interaction data and mission-critical trajectories separately, thereby improving experience utilization during training. The study also formulated the state space, action space, and reward function for target exploration, obstacle avoidance, energy consumption, and velocity maintenance under limited perception conditions. The experiments evaluated the proposed algorithm in a two-dimensional tracking scenario with moving targets, multiple unmanned aerial vehicles, and dynamic and static obstacles. The proposed method was compared with Deep Deterministic Policy Gradient and Multi-Agent Deep Deterministic Policy Gradient using collision rate, capture rate, capture time, and capture distance as evaluation metrics. The results showed that the proposed algorithm improved convergence behavior and average reward, reduced the collision rate by 83% compared with Deep Deterministic Policy Gradient at the first environmental level, and maintained competitive performance in capture rate, capture time, and path efficiency.
Qinglin Han, Hongmei Wang· International Conference on...· 0 citations
Multi-UAV Cooperative Target Search (MCTS) is a critical task in low-altitude sensing applications, requiring agents to efficiently explore unknown environments under complex constraints. However, traditional search methods are mostly unscalable and perform poorly in dynamic multi-UAV environments. As a promising alternative, Reinforcement Learning (RL) has emerged to overcome these limitations by enabling agents to learn adaptive policies directly from environmental interactions. A key limitation is that current RL methods lack efficient exploration, which is a critical bottleneck preventing UAVs from finding more targets. To address this limitation, we propose a novel method named AEQMIX, which integrates trajectory entropy maximization into QMIX, an advanced Multi-Agent Reinforcement Learning (MARL) method, to encourage efficient exploration. We formulate the MCTS problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and design a multi-objective reward function. To mitigate the intractability of density estimation in high-dimensional spaces, we employ a nonparametric particle-based entropy estimator to quantify the spatial diversity of UAV trajectories. This entropy estimate is utilized as an intrinsic reward, incentivizing agents to maximize the distance between their trajectories and those of their neighbors. Extensive simulations demonstrate that AEQMIX significantly outperforms baseline reinforcement learning and traditional optimization methods in terms of search rate, coverage efficiency, and collision avoidance. Compared with DNQMIX, AEQMIX improves the search rate and coverage rate by 9.52% and 11.54%, respectively, while reducing the average collision count by 70.59% in the (40 × 40) environment.
Unmanned aerial vehicles (UAVs) offer several advantages, including high mobility, flexible deployment, low cost, and strong adaptability to complex environments, making them highly promising for applications such as disaster search and rescue, environmental monitoring, inspection, and reconnaissance. For target exploration tasks in unknown environments, multiUAV systems can expand the search area, improve exploration efficiency, and enhance the robustness of task execution through cooperation, which makes this problem of significant research interest. However, such tasks still face several challenges, including partial observability of environmental information, complex cooperative decision-making, and difficulties in credit assignment among multiple UAVs. Reinforcement learning is capable of learning decision-making policies autonomously through interaction with the environment, providing a new perspective for solving cooperative exploration problems in complex environments. To address these issues, we propose a cooperative decision-making method for multi-UAV target exploration. By incorporating target-related information, the proposed method enhances the cooperative exploration capability of UAVs in unknown environments, while a tailored reward design is adopted to improve the coordination efficiency of multiple UAVs. Experimental results show that the proposed method exhibits strong adaptability to different team sizes and sensor configurations, learns effective cooperative behaviors, and outperforms classical exploration methods across multiple performance metrics, thereby demonstrating its effectiveness in multi-UAV target exploration tasks.
Batuo Zhang, Lei Liu, Zhongmin Yan et al.· Fall Joint Computer Conferen...· 0 citations
A multi-agent deep reinforcement learning framework that addresses issues through coordinated exploration, demonstration exploitation, safe curriculum scheduling, and structure-aware generalisation is proposed, demonstrating strong performance in collaboration success rate, navigation robustness, zero-shot cross-scenario generalisation, and dynamic environment adaptability.
Yuhuang Su, Nabil Aouf· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.