Jul 2026· Fall Joint Computer Conference· pp. 225-232· 0 citations· 30 references
Abstract
Unmanned aerial vehicles (UAVs) offer several advantages, including high mobility, flexible deployment, low cost, and strong adaptability to complex environments, making them highly promising for applications such as disaster search and rescue, environmental monitoring, inspection, and reconnaissance. For target exploration tasks in unknown environments, multiUAV systems can expand the search area, improve exploration efficiency, and enhance the robustness of task execution through cooperation, which makes this problem of significant research interest. However, such tasks still face several challenges, including partial observability of environmental information, complex cooperative decision-making, and difficulties in credit assignment among multiple UAVs. Reinforcement learning is capable of learning decision-making policies autonomously through interaction with the environment, providing a new perspective for solving cooperative exploration problems in complex environments. To address these issues, we propose a cooperative decision-making method for multi-UAV target exploration. By incorporating target-related information, the proposed method enhances the cooperative exploration capability of UAVs in unknown environments, while a tailored reward design is adopted to improve the coordination efficiency of multiple UAVs. Experimental results show that the proposed method exhibits strong adaptability to different team sizes and sensor configurations, learns effective cooperative behaviors, and outperforms classical exploration methods across multiple performance metrics, thereby demonstrating its effectiveness in multi-UAV target exploration tasks.
Multi-UAV Cooperative Target Search (MCTS) is a critical task in low-altitude sensing applications, requiring agents to efficiently explore unknown environments under complex constraints. However, traditional search methods are mostly unscalable and perform poorly in dynamic multi-UAV environments. As a promising alternative, Reinforcement Learning (RL) has emerged to overcome these limitations by enabling agents to learn adaptive policies directly from environmental interactions. A key limitation is that current RL methods lack efficient exploration, which is a critical bottleneck preventing UAVs from finding more targets. To address this limitation, we propose a novel method named AEQMIX, which integrates trajectory entropy maximization into QMIX, an advanced Multi-Agent Reinforcement Learning (MARL) method, to encourage efficient exploration. We formulate the MCTS problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and design a multi-objective reward function. To mitigate the intractability of density estimation in high-dimensional spaces, we employ a nonparametric particle-based entropy estimator to quantify the spatial diversity of UAV trajectories. This entropy estimate is utilized as an intrinsic reward, incentivizing agents to maximize the distance between their trajectories and those of their neighbors. Extensive simulations demonstrate that AEQMIX significantly outperforms baseline reinforcement learning and traditional optimization methods in terms of search rate, coverage efficiency, and collision avoidance. Compared with DNQMIX, AEQMIX improves the search rate and coverage rate by 9.52% and 11.54%, respectively, while reducing the average collision count by 70.59% in the (40 × 40) environment.
Distributed multi-UAV systems play an important role in applications such as search and rescue, disaster response, environmental monitoring, and autonomous reconnaissance. These tasks often require multiple UAVs to coordinate navigation and sensing so as to improve efficiency and expand useful environment coverage. However, in goal-directed collaborative exploration, it remains difficult to balance rapid target reaching with effective exploration of unknown areas, especially when redundant sensing and local spatial competition must also be considered. To address this challenge, we propose SPICE, a communication framework for goal-directed collaborative exploration under a frontier-graph action abstraction. The proposed framework improves coordination by learning more informative and interpretable communication among agents and by encouraging behaviors that reduce local overlap during navigation. Experimental results show that SPICE achieves a better balance between exploration quality and coordination efficiency than representative value-based baselines, yielding higher coverage and lower observation redundancy while maintaining competitive target-reaching performance.
Cheng-Lin Tang, Lei Liu, Xudong Lu et al.· Fall Joint Computer Conferen...· 0 citations
In recent years, Unmanned Aerial Vehicles (UAVs) have gradually been widely used in various fields such as regional search and disaster relief, and the development of related technologies has also experienced unprecedented growth. Compared to individual UAVs, the collaborative execution of tasks by UAV swarms has more advantages, but it is often difficult to achieve fast and accurate autonomous navigation and obstacle avoidance capabilities in unknown complex obstacle environments due to limitations in computing load and inter-UAV communication capabilities. To solve this problem, a hierarchical navigation decision-making framework for non-communication UAVs in unknown environments is proposed in this paper, which decomposes the UAV navigation planning task into an upper-layer global planning module and a lower-layer autonomous navigation and obstacle avoidance module. For the lower-layer module, an enhanced hybrid feature extraction network is designed, accompanied by a dual-stage training strategy that integrates traditional optimization methods with reinforcement learning. The upper-layer module incorporates a hybrid control strategy combining conventional search methods. Based on the ROS framework, simulation experiments for UAV swarm navigation were systematically conducted. The experimental results demonstrate that the proposed algorithm achieves autonomous navigation decision-making for multiple UAVs in complex unknown obstacle environments without relying on inter-UAV communication, showing significant advantages compared to existing approaches.
This paper analyzes the factors affecting communication interactions between UAVs and proposes a bidding-based grouping method to eliminate ineffective communication interactions, and introduces a network simplification algorithm based on reducing the number of triangular network topologies to optimize the communication network structure.
Wei-Xing Xia, Peng Chen, Fei-Fei Song et al.· Drones· 0 citations
A cooperative guidance law based on the experience-guided multi-agent proximal policy optimization (E-MAPPO) algorithm is proposed for multiple unmanned aerial vehicles (UAVs) to track dynamic points of interest in civilian applications and results indicate that the proposed method generalizes well to different types of maneuvering targets.
Hao Xiong, Minghu Tan, Xiaoyu Liu et al.· Drones· 0 citations
Experimental results demonstrate that the Improved Experience Replay Multi-Agent Deep Deterministic Policy Gradient algorithm outperforms other comparison algorithms in terms of convergence speed, training stability, and path-planning performance.
Long Wen, Hui Tan, Yuxi Liu et al.· Italian National Conference...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.