Jul 2026· IEEE Transactions on Neural Networks and Learning Systems· Vol PP, pp. 1-15· 0 citations
Medicine
TL;DR
This work proposes a distributed Multiagent Energy-saving Autonomous Exploration System (MEAES), and devise the consumption-exploration-balanced training framework (CEBF), which guides agents to transform from lazy exploration to energy-saving exploration strategies through dynamic reward shaping.
Abstract
Multiagent autonomous exploration in unknown environments is both meaningful and challenging. Due to the constraint of a partially observable environment, the collaboration among agents is often inadequate, leading to increased energy consumption. Worse still, a decrease in overall exploration performance may occur due to a single agent failure. To address these issues, we propose a distributed Multiagent Energy-saving Autonomous Exploration System (MEAES) based on reinforcement learning. To accurately evaluate the regional complexity of different branches and further enhance the long-term decision-making capabilities of agents, we introduce the dual-scale clustered observation (DSCO) module. The DSCO generates fine-grained representations based on graph modeling, enabling better characterization of both global and long-term exploration values. Furthermore, we propose an energy-saving action (EA) mechanism, which mitigates redundant exploration and reduces energy consumption by selective waiting actions and independent exploration strategies. Finally, we devise the consumption-exploration-balanced training framework (CEBF), which guides agents to transform from lazy exploration to energy-saving exploration strategies through dynamic reward shaping. Extensive experiments validate the effectiveness of MEAES, demonstrating effective zero-shot transfer performance across unseen environments.
Multi-UAV Cooperative Target Search (MCTS) is a critical task in low-altitude sensing applications, requiring agents to efficiently explore unknown environments under complex constraints. However, traditional search methods are mostly unscalable and perform poorly in dynamic multi-UAV environments. As a promising alternative, Reinforcement Learning (RL) has emerged to overcome these limitations by enabling agents to learn adaptive policies directly from environmental interactions. A key limitation is that current RL methods lack efficient exploration, which is a critical bottleneck preventing UAVs from finding more targets. To address this limitation, we propose a novel method named AEQMIX, which integrates trajectory entropy maximization into QMIX, an advanced Multi-Agent Reinforcement Learning (MARL) method, to encourage efficient exploration. We formulate the MCTS problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and design a multi-objective reward function. To mitigate the intractability of density estimation in high-dimensional spaces, we employ a nonparametric particle-based entropy estimator to quantify the spatial diversity of UAV trajectories. This entropy estimate is utilized as an intrinsic reward, incentivizing agents to maximize the distance between their trajectories and those of their neighbors. Extensive simulations demonstrate that AEQMIX significantly outperforms baseline reinforcement learning and traditional optimization methods in terms of search rate, coverage efficiency, and collision avoidance. Compared with DNQMIX, AEQMIX improves the search rate and coverage rate by 9.52% and 11.54%, respectively, while reducing the average collision count by 70.59% in the (40 × 40) environment.
The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.
X.-H. Fang, K. Chen, Cheng-Hao Ren et al.· Advanced Electromagnetics· 0 citations
This work introduces a novel MARL framework, Multi-Agent Divergence Policy Optimization (MADPO) with Mutual Policy Divergence Maximization (Mutual PDM), and proposes a new extension of CCS divergence for measuring policy divergence of more than two agents, the Generalized Conditional Cauchy-Schwarz (GCCS) divergence.
Haowen Dou, Lujuan Dang, Mingfei Lu et al.· IEEE Transactions on Pattern...· 0 citations
A value-aware extension of Multi-Agent Observation Sharing under Communication Dropout to patch communication gaps is proposed; it is referred to as Value-Aware MARO and dynamically weighting the predictor's loss function using advantage estimates derived from the underlying actor-critic architecture.
K. D. Kafadar, Eren Özaltun, M. E. Şanlı et al.· arXiv.org· 0 citations
Action Generation with Topology Awareness (AGTA), a topology-aware sequential decision-making framework in MARL that integrates inter-agent correlation modeling with topology-guided decision-order optimization, and outperforms the state-of-the-art counterparts.
Kun Hu, Shanghua Wen, Wendi Wu et al.· Mathematics· 0 citations
The findings demonstrate MARL’s promise in solving navigation problems efficiently and provide concrete recommendations for tuning training parameters and network structures to enhance performance and robustness.
Stanislav Safranek, Brian M. Kirk· International Journal of Inn...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.