Skip to content

Safe Decision-Making via Adaptive Causal Representation for Autonomous Driving.

Jul 2026 · IEEE Transactions on Neural Networks and Learning Systems · Vol PP, pp. 1-14 · 0 citations
Medicine

TL;DR

This work proposes a safe offline-to-online decision-making framework with adaptive causal representation, an adaptive causal transformer (AC-Transformer), which learns a causal representation from offline driving trajectories that is adaptive as the deployment distribution evolves.

Abstract

Offline reinforcement learning (RL) is promising for autonomous driving, but as deployment conditions drift away from the offline training distribution, policies may encounter out-of-distribution (OOD) scenarios, such as unseen road geometries and diverse driving behaviors, rendering offline-learned decisions unreliable. To address this issue, we propose a safe offline-to-online decision-making framework with adaptive causal representation. At its core is an adaptive causal transformer (AC-Transformer), which learns a causal representation from offline driving trajectories. The representation is causal because it captures how states and actions influence long-horizon reward and safety outcomes under deployment shift. As these influences evolve over sequential interactions, we model them jointly over states, actions, rewards, and costs. Then, bisimulation regularization is introduced to further organize the learned representation into a consequence-consistent latent structure, so that it reflects long-horizon consequences rather than superficial traffic patterns. During online deployment, a phase-adaptive objective (PAC) is designed to progressively refine the learned representation, making it adaptive as the deployment distribution evolves. We evaluate the proposed method in five distinct driving scenarios against eight representative baselines. The results demonstrate stronger generalization across diverse road structures and stochastic driving styles while preserving a favorable balance between safety and efficiency.

View source

Similar papers

Preprint Sep 2026

From Proxy Learning to Driving Decisions: A Transfer-Based Framework for Evaluating Future-Aware Autonomous Driving Planners

Future-aware representations and world models are increasingly used in proposal-based autonomous-driving planners to improve trajectory selection. However, improvements in proxy objectives or restricted subsets are often interpreted as planning gains without verifying proposal ordering, selected trajectories, full-scal...

Yi-Kai Wu · 0 citations
Preprint Aug 2026

Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving

Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is...

Dong Hu, Chao Huang, Carman K. M. Lee et al. · 0 citations
Preprint Aug 2026

Knowledge-Data-Dual-Driven Reinforcement Learning for Autonomous Vehicle Control in Mixed Traffic

In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver behaviors, limiting the proactive reasoning capabilities. Sec...

Jie Fang, Wei Zheng, Mengyun Xu et al. · 0 citations
Conference Aug 2026

CoPL: Counterfactual Offline Policy Learning via Large Language Models

Offline reinforcement learning is limited by the coverage of the offline dataset, which makes it difficult for policies to generalize to unseen goals and behaviors. We introduce CoPL, a framework for counterfactual offline policy learning via large language models. Unlike prior relabeling approaches that modify only in...

Maosen Zeng, Yunan Liu · 0 citations
Open access Aug 2026

Dual-critic constrained deceptive Q-learning for deployment-time policy protection

The Dual-Critic Constrained Deceptive Q-Learning (DCD-Q) method is proposed, a deployment-time trajectory protection framework that aims to reduce the information leaked by released trajectories while preserving acceptable task performance.

Guang-Yu Pan, Bo Hou, Yao Chen et al. · 0 citations
Preprint Aug 2026

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to...

Xincong Hu, Lei Ou, Maosen Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.