An algorithm-agnostic Distributionally Robust MARL framework integrating an adaptive Contextual-Bandit Worst-Case Estimator (CB-WCE) co-evolves with the traffic controllers by dynamically generating adversarial demand mixtures during training, demonstrating the framework's scalability and potential for resilient real-world urban deployment.
Abstract
Multi-agent reinforcement learning (MARL) has emerged as a promising approach for traffic signal control. However, standard MARL policies typically optimize for expected returns under nominal conditions, leaving them highly vulnerable to spatial-temporal demand shifts and catastrophic congestion under adverse scenarios. To address this critical limitation, this paper proposes an algorithm-agnostic Distributionally Robust (DR) MARL framework integrating an adaptive Contextual-Bandit Worst-Case Estimator (CB-WCE). Operating on a slower timescale, the CB-WCE co-evolves with the traffic controllers by dynamically generating adversarial demand mixtures during training. This steers the learning process to fortify policies against bottleneck scenarios without requiring modifications to the underlying MARL architectures. The framework is evaluated across value-based, actor-critic, and policy-gradient methods on both a synthetic 5x5 grid and a heterogeneous Monaco City network. Empirical results demonstrate that the DR framework prevents unbounded queue growth and profoundly enhances both worst-case robustness and average-case efficiency. Notably, for the Proximal Policy Optimization (PPO) architecture in the Monaco environment, on average, robust retraining reduced the worst-case queue length by 74.39% and improved the average-case network-wide queue length by 75.45%. Furthermore, the retrained policies exhibit strong zero-shot generalization to unseen traffic distributions, highlighting the framework's scalability and potential for resilient real-world urban deployment.
DRIQN is proposed to integrate Distributionally Robust Optimization (DRO) with implicit quantile networks to optimize worst-case performance under natural environmental conditions and incorporates heterogeneous noise sources and target robustness-critical scenarios.
Zhao-Fan Zhang, Minghao Yang, Si-Hong Xie et al.· 0 citations
This study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO), which successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshear escape and future competency-based flight training.
Multi-unmanned surface vehicle (USV) pursuit–evasion missions in maritime environments presents significant challenges due to dynamic ship populations, high-dimensional observations, and the gap between idealised simulations and real-world maritime physics. To address these challenges, we propose a Credit-Aware Multi-Agent Reinforcement Learning (CA-MARL) framework for multi-USV pursuit–evasion. The framework features two key innovations: a Residual Self-Attention module that adapts to varying fleet sizes through permutation-invariant attention, and a Mixed Credit Assignment module that enhances centralised value estimation with decentralised branches. Moreover, to bridge the simulation-to-reality gap, we develop a high-fidelity 3D virtual platform using Unity3D that incorporates maritime factors, such as hydrodynamics and wave disturbances, which are typically overlooked in USV simulations but critical for maritime operations. Experiments demonstrate that our method achieves superior coordination, sample efficiency, and policy robustness compared to existing baselines, providing a credible foundation for deploying MARL policies in realistic multi-USV scenarios.
A value-aware extension of Multi-Agent Observation Sharing under Communication Dropout to patch communication gaps is proposed; it is referred to as Value-Aware MARO and dynamically weighting the predictor's loss function using advantage estimates derived from the underlying actor-critic architecture.
K. D. Kafadar, Eren Özaltun, M. E. Şanlı et al.· arXiv.org· 0 citations
Accurate multiagent trajectory forecasting is paramount for the safety of autonomous driving systems, yet existing methods frequently struggle to balance high predictive fidelity with the computational efficiency required for real-time deployment. This study proposes a rule-guided lightweight framework (RuLiF), a novel approach designed to address this trade-off by explicitly integrating traffic-rule priors into the representation learning process. The methodology centers on a rule-guided symmetric fusion transformer that employs a bidirectional attention mechanism to unify dynamic agent interactions and static map topologies. This fusion is dynamically modulated by an 18-dimensional rule feature vector that strictly encodes kinematic states, collision risks, and lane constraints. To guarantee kinematic feasibility, the framework utilizes a Bézier-parameterized decoder that generates smooth continuous curves, complemented by a training-time constraint rectification strategy. This strategy applies projected gradient descent to the top-three confident modes during training to enforce safety norms—such as yielding protocols and safe following distances—without incurring any additional inference latency. Extensive experiments on two large-scale motion-forecasting benchmarks validate the effectiveness of RuLiF under both short- and long-horizon prediction settings. On the short-horizon benchmark, RuLiF achieves a minimum final displacement error of 0.95 m and a miss rate of 0.08 with only 1.9 million parameters. On the more challenging long-horizon benchmark, the model obtains a minimum average displacement error of 0.78 m, a minimum final displacement error of 1.45 m, and a miss rate of 0.20 while using 2.6 million parameters and only 34% of the floating-point operations (FLOPs) required by the most computationally intensive compared model. Furthermore, the rectification module significantly enhances social compliance, reducing the time-to-collision critical rate to 1.8% and lane violations to a mere 0.9%, demonstrating that RuLiF successfully harmonizes geometric reasoning with behavioral compliance for resource-constrained onboard applications.
Shangguan Wei, Mingzhe Huang, Linguo Chai et al.· Journal of Transportation En...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.