Aug 2026· 2026 3rd International Conference on Intelligent Systems and Robotics (CISR)· pp. 1-5· 0 citations· 10 references
Abstract
Unmanned surface and underwater vehicles face challenges in autonomously learning cooperative encirclement for high-value targets under partial observability, intermittent communication, and collision risks. This paper proposes a heterogeneous multi-agent reinforcement learning framework with safe decisionmaking. The framework includes a Dec-POMDP-based heterogeneous architecture that unifies state representation, refines platform attributes, and differentiates action constraints, enabling distinct platforms to learn specialized policies within a shared model. A weighted composite reward function addresses target distance, formation geometry, safety, coordination consistency, action smoothness, and task completion, with an implicit coordination reward embedded to balance local control and sustained cooperation under limited communication. Building on MATD3, this paper introduces Feasible Safe Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (FS-MATD3), which incorporates a dual-safety gateway through training-stage cost and execution-stage correction, alongside a three-stage training strategy (normal, dropout, anomaly) and observation masking to improve robustness against link and node failures. Experiments in a simulated marine environment show that the pursuers' cross-domain success rate converges to a stable level and maintains that stability throughout later training; the evader's maneuverable space shrinks progressively and cross-run variance declines over trials. The method demonstrates consistent reward convergence and task completion reliability across runs, providing effective decision support for safe cooperative encirclement by heterogeneous unmanned platforms.
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative navigation with time-varying topology. To address the critical challenge of ensuring safety in environments with sensing constraints, a safety-decoupled mechanism is introduced through a Control Barrier-Like Funct...
Sizhe Xiao, Li-Jing Dong, Rui-Ting Bai et al.· 0 citations
MA-PMBRL, a novel Multi-Agent Pessimistic Model-Based Reinforcement Learning framework for CAVs, incorporating a max-min optimization approach to enhance robustness and decision-making is proposed, demonstrating that the proposed framework represents a significant step toward scalable, efficient, and reliable multi-age...
Ruo-Qi Wen, Rong-Peng Li, Xing Xu et al.· IEEE Transactions on Mobile...· 1 citation
The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.
X.-H. Fang, K. Chen, Cheng-Hao Ren et al.· Advanced Electromagnetics· 0 citations
Cooperative navigation of multiple unmanned aerial vehicles (UAVs) in disaster search-and-rescue scenarios is challenging due to dense obstacles, partial observability, and strong inter-agent coupling, which often result in path conflicts, collision risks, and limited policy generalization. To address these challenges,...
Li Tan, Hai-Xia Zhao, Jia-Qin Chai et al.· Unmanned Systems· 0 citations
This paper proposes SafeGPT, a hierarchical framework that integrates generative pretrained transformer (GPT)-based large language models (LLM) with safe reinforcement learning (safe-RL). SafeGPT addresses the large-scale random multi-point tour problem (RMPT) for multiple unmanned aerial vehicles (UAVs). The target pl...
Hyojun Ahn, Seungcheol Oh, Gyusun Kim et al.· IEEE Transactions on Network...· 0 citations
A novel deep reinforcement learning framework that enables safe and efficient navigation in such communication-free, signal-free, and lane-free intersection environments while meeting stringent safety requirements for practical deployment is proposed.
Ruihao Zeng, M. Ramezani· IEEE Transactions on Cyberne...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.