Temporal coordination aware reinforcement learning for multi-agent UAV navigation in dynamic environments
TL;DR
T-CARE is introduced, a hybrid learning-heuristic multi-UAV coordination framework that integrates zero-shot constrained action selection with priority-aware temporal reservations and achieves 100% success, 0% collision rate, and no observed persistent starvation or deadlock.
Abstract
Multi-agent UAV navigation in cluttered environments is challenged by deadlock, starvation-like persistent yielding, oscillation, and unsafe congestion, particularly in narrow passages and dynamically emerging bottlenecks. Existing approaches often rely on static planning assumptions, reciprocal compliance, or learning-based policies in which temporal ordering and fairness are not explicitly represented. This paper introduces T-CARE (Temporal Coordination-Aware Reinforcement Learning), a hybrid learning-heuristic multi-UAV coordination framework that integrates zero-shot constrained action selection with priority-aware temporal reservations. T-CARE extends the prior ZSE-CRL action-selection backbone by adding: (i) runtime spatiotemporal reservations, (ii) priority aging and wait-time accounting designed to reduce persistent yielding, (iii) temporary-goal reassignment for blocked or congested motion, (iv) stagnation detection with escape and recovery, and (v) geometry-agnostic bottleneck discovery with shared-corridor reuse. We evaluate T-CARE in simulation using three-swarm adversarial congestion stress tests, and ten-swarm reconstructed 3D urban/suburban environments, with each ten-swarm scenario containing one leader and three followers per swarm for a total of 40 UAVs, under strict zero-shot deployment. Across the tested scenarios, T-CARE achieved 100% success, 0% collision rate, and no observed persistent starvation or deadlock, while learning-only, reactive, and coordination-reduced baselines exhibited failures under the same evaluation protocol. These results support T-CARE as an empirically effective coordination architecture for simulated multi-UAV traffic, while formal convergence-time guarantees, communication-delay robustness, and real-world flight validation remain future work.