Skip to content
Open access

Multi-Agent Uncertainty-Aware Pessimistic Model-Based Reinforcement Learning for Connected Autonomous Vehicles

Mar 2025 · IEEE Transactions on Mobile Computing · Vol 25, pp. 16522-16539 · 1 citation · 64 references
Computer Science

TL;DR

MA-PMBRL, a novel Multi-Agent Pessimistic Model-Based Reinforcement Learning framework for CAVs, incorporating a max-min optimization approach to enhance robustness and decision-making is proposed, demonstrating that the proposed framework represents a significant step toward scalable, efficient, and reliable multi-agent decision-making for CAVs.

Abstract

Deep Reinforcement Learning (DRL) holds significant promise for achieving human-like Autonomous Vehicle (AV) capabilities, but suffers from low sample efficiency and challenges in reward design. Model-Based Reinforcement Learning (MBRL) offers improved sample efficiency and generalizability compared to Model-Free Reinforcement Learning (MFRL) in various multi-agent decision-making scenarios. Nevertheless, MBRL faces critical difficulties in estimating uncertainty during the model learning phase, thereby limiting its scalability and applicability in real-world scenarios. Additionally, most studies on Connected Autonomous Vehicles (CAVs) focus on single-agent decision-making. In contrast, existing multi-agent MBRL solutions lack computationally tractable algorithms with Probably Approximately Correct (PAC) guarantees, a crucial factor for ensuring policy reliability with limited training data. To address these challenges, we propose MA-PMBRL, a novel Multi-Agent Pessimistic Model-Based Reinforcement Learning framework for CAVs, incorporating a max-min optimization approach to enhance robustness and decision-making. To mitigate the inherent subjectivity of uncertainty estimation in MBRL and avoid incurring catastrophic failures in AV, MA-PMBRL employs a pessimistic optimization framework combined with Projected Gradient Descent (PGD) for both model and policy learning. MA-PMBRL also employs general function approximations under partial dataset coverage to enhance learning efficiency and system-level performance. By bounding the suboptimality of the resulting policy under mild theoretical assumptions, we successfully establish PAC guarantees for MA-PMBRL, demonstrating that the proposed framework represents a significant step toward scalable, efficient, and reliable multi-agent decision-making for CAVs.

Read PDF

Similar papers

Review Open access Aug 2026

Advances in Multi-Agent Deep Reinforcement Learning: Methods with Applications and Challenges

This paper presents a narrative survey of recent developments in MARL and examines research directions centred on centralised training with decentralised execution (CTDE), value decomposition, learned communication, graph-based methods, and model-based learning.

Abdur Rakib, Khoa Phung, Marco Pérez Hernández et al. · 0 citations
Conference Aug 2026

Risk-Informed Multi-Agent Reinforcement Learning for Embedded Systems on Resource-Constrained Hardware

Learning-enabled control systems increasingly rely on multi-agent reinforcement learning to operate in uncertain and interactive environments. While risk-aware decision-making has been shown to improve safety and robustness, deploying such algorithms on resource-constrained embedded platforms remains a significant chal...

Lachlan Talento, Kyle Pham, Bhaskar Ramasubramanian · 0 citations
Open access Aug 2026

Efficient Exploration-Enabled Multi-Agent Reinforcement Learning for Multi-UAV Cooperative Target Search

Multi-UAV Cooperative Target Search (MCTS) is a critical task in low-altitude sensing applications, requiring agents to efficiently explore unknown environments under complex constraints. However, traditional search methods are mostly unscalable and perform poorly in dynamic multi-UAV environments. As a promising alter...

Peng Chen, Tian-Xu Li, Wei-Xing Xia et al. · 0 citations
Conference Aug 2026

Heterogeneous Multi-Agent Autonomous Learning and Safe Cooperative Decision-Making

Unmanned surface and underwater vehicles face challenges in autonomously learning cooperative encirclement for high-value targets under partial observability, intermittent communication, and collision risks. This paper proposes a heterogeneous multi-agent reinforcement learning framework with safe decisionmaking. The f...

Jiang-Li Cao, Chao Liu, Guo-Ping Zhang · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations
Aug 2026

Deep reinforcement learning–based safe path planning for leader–follower robots

This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.

Ehsan Kazemi Tameh, Mohammadreza Estarki, Saeed Khodaygan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.