Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 18 references
Abstract
Swarm unmanned surface vehicle (USV) has become a promising solution for maritime security and defense activities. However, the complexity of continuous multi-target hunting presents significant challenges for multi-agent coordination in water environments. This study proposes an XAI-driven multi-agent reinforcement learning (MARL) framework for swarm USVs to perform continuous multi-target hunting. The framework integrates multi-agent deep reinforcement learning (basic MAPPO and MAPPO-LSTM) for cooperative hunting policies with explainable AI (XAI) mechanisms to provide interpretable insights into swarm behaviors and decision strategies. Experimental results in a 3D simulation platform environment demonstrate that MAPPO-LSTM achieves superior interception performance, reducing mean time-to-capture and improving trajectory smoothness compared to the baseline MAPPO. Furthermore, the proposed XAI pipeline combines distance-based importance attribution, influence graph analysis, strategic clustering, and temporal dynamics evaluation to explain agent contributions, inter-agent dependencies, emergent strategies, and efficiency patterns. By providing multi-level interpretability, the framework enhances transparency, trust, and deployment readiness of swarm USVs in dynamic maritime defense scenarios.
The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.
X.-H. Fang, K. Chen, Cheng-Hao Ren et al.· Advanced Electromagnetics· 0 citations
Multi-UAV Cooperative Target Search (MCTS) is a critical task in low-altitude sensing applications, requiring agents to efficiently explore unknown environments under complex constraints. However, traditional search methods are mostly unscalable and perform poorly in dynamic multi-UAV environments. As a promising alternative, Reinforcement Learning (RL) has emerged to overcome these limitations by enabling agents to learn adaptive policies directly from environmental interactions. A key limitation is that current RL methods lack efficient exploration, which is a critical bottleneck preventing UAVs from finding more targets. To address this limitation, we propose a novel method named AEQMIX, which integrates trajectory entropy maximization into QMIX, an advanced Multi-Agent Reinforcement Learning (MARL) method, to encourage efficient exploration. We formulate the MCTS problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and design a multi-objective reward function. To mitigate the intractability of density estimation in high-dimensional spaces, we employ a nonparametric particle-based entropy estimator to quantify the spatial diversity of UAV trajectories. This entropy estimate is utilized as an intrinsic reward, incentivizing agents to maximize the distance between their trajectories and those of their neighbors. Extensive simulations demonstrate that AEQMIX significantly outperforms baseline reinforcement learning and traditional optimization methods in terms of search rate, coverage efficiency, and collision avoidance. Compared with DNQMIX, AEQMIX improves the search rate and coverage rate by 9.52% and 11.54%, respectively, while reducing the average collision count by 70.59% in the (40 × 40) environment.
With the increasing deployment of multi-unmanned aerial vehicle (multi-UAV) systems in dynamic environments, the problem of efficient cooperative path planning has emerged as a critical challenge requiring urgent solutions. To address this issue, this paper proposes a novel joint optimization framework, named spatio-temporal attention-based multi-agent deep deterministic policy gradient (STA-MADDPG). Rather than proposing a new reinforcement learning algorithm in the strict sense, this work integrates advanced spatial-temporal feature extraction with heuristic gradient guidance. First, a cascaded architecture combining multi-head attention and Long Short-Term Memory (LSTM) networks is utilized to extract key local and temporal features, thereby mitigating the dimensionality curse in dense multi-agent observations. Second, an improved dynamic artificial potential field (DAPF) is integrated into the reinforcement learning framework as a state augmentation mechanism, providing heuristic guidance vectors that accelerate convergence and improve obstacle avoidance. Furthermore, to balance computational complexity and adaptive behavior, a rule-based hierarchical formation strategy is designed. The framework maps predefined formations (elliptical, chain, or wedge) to specific environment categories, while the underlying MARL policy governs the dynamic trajectory planning and topology maintenance. Finally, rigorous comparative and ablation experiments are conducted to evaluate path length, search time, and relative position errors. Statistical analysis demonstrates the effectiveness of the proposed framework, achieving up to a 67.3% reduction in search time and a 91.56% search success rate compared with standard MARL baselines in complex environments.
Multi-unmanned surface vehicle (USV) pursuit–evasion missions in maritime environments presents significant challenges due to dynamic ship populations, high-dimensional observations, and the gap between idealised simulations and real-world maritime physics. To address these challenges, we propose a Credit-Aware Multi-Agent Reinforcement Learning (CA-MARL) framework for multi-USV pursuit–evasion. The framework features two key innovations: a Residual Self-Attention module that adapts to varying fleet sizes through permutation-invariant attention, and a Mixed Credit Assignment module that enhances centralised value estimation with decentralised branches. Moreover, to bridge the simulation-to-reality gap, we develop a high-fidelity 3D virtual platform using Unity3D that incorporates maritime factors, such as hydrodynamics and wave disturbances, which are typically overlooked in USV simulations but critical for maritime operations. Experiments demonstrate that our method achieves superior coordination, sample efficiency, and policy robustness compared to existing baselines, providing a credible foundation for deploying MARL policies in realistic multi-USV scenarios.
An Attention-based Multi-Agent Deep Deterministic Policy Gradient algorithm was developed for cooperative multi-unmanned aerial vehicle target tracking in dynamic environments. The study addressed information redundancy and association weight allocation between individual unmanned aerial vehicles and the swarm during cooperative decision-making. To improve information selection, the proposed algorithm introduced a centralized critic network with a multi-head attention mechanism to evaluate the contributions of other agents at each time step. Meanwhile, the study designed a dual-buffer experience management architecture composed of a recent interaction memory and a mission outcome memory. This architecture stored recent interaction data and mission-critical trajectories separately, thereby improving experience utilization during training. The study also formulated the state space, action space, and reward function for target exploration, obstacle avoidance, energy consumption, and velocity maintenance under limited perception conditions. The experiments evaluated the proposed algorithm in a two-dimensional tracking scenario with moving targets, multiple unmanned aerial vehicles, and dynamic and static obstacles. The proposed method was compared with Deep Deterministic Policy Gradient and Multi-Agent Deep Deterministic Policy Gradient using collision rate, capture rate, capture time, and capture distance as evaluation metrics. The results showed that the proposed algorithm improved convergence behavior and average reward, reduced the collision rate by 83% compared with Deep Deterministic Policy Gradient at the first environmental level, and maintained competitive performance in capture rate, capture time, and path efficiency.
Qinglin Han, Hongmei Wang· International Conference on...· 0 citations
This paper presents a narrative survey of recent developments in MARL and examines research directions centred on centralised training with decentralised execution (CTDE), value decomposition, learned communication, graph-based methods, and model-based learning.
Abdur Rakib, K. Phung, Marco Pérez Hernández et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.