Jul 2026· IEEE Transactions on Neural Networks and Learning Systems· Vol PP, pp. 1-14· 0 citations
Medicine
TL;DR
This work proposes a safe offline-to-online decision-making framework with adaptive causal representation, an adaptive causal transformer (AC-Transformer), which learns a causal representation from offline driving trajectories that is adaptive as the deployment distribution evolves.
Abstract
Offline reinforcement learning (RL) is promising for autonomous driving, but as deployment conditions drift away from the offline training distribution, policies may encounter out-of-distribution (OOD) scenarios, such as unseen road geometries and diverse driving behaviors, rendering offline-learned decisions unreliable. To address this issue, we propose a safe offline-to-online decision-making framework with adaptive causal representation. At its core is an adaptive causal transformer (AC-Transformer), which learns a causal representation from offline driving trajectories. The representation is causal because it captures how states and actions influence long-horizon reward and safety outcomes under deployment shift. As these influences evolve over sequential interactions, we model them jointly over states, actions, rewards, and costs. Then, bisimulation regularization is introduced to further organize the learned representation into a consequence-consistent latent structure, so that it reflects long-horizon consequences rather than superficial traffic patterns. During online deployment, a phase-adaptive objective (PAC) is designed to progressively refine the learned representation, making it adaptive as the deployment distribution evolves. We evaluate the proposed method in five distinct driving scenarios against eight representative baselines. The results demonstrate stronger generalization across diverse road structures and stochastic driving styles while preserving a favorable balance between safety and efficiency.
Future-aware representations and world models are increasingly used in proposal-based autonomous-driving planners to improve trajectory selection. However, improvements in proxy objectives or restricted subsets are often interpreted as planning gains without verifying proposal ordering, selected trajectories, full-scal...
Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is...
Dong Hu, Chao Huang, Carman K. M. Lee et al.· 0 citations
In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver behaviors, limiting the proactive reasoning capabilities. Sec...
Jie Fang, Wei Zheng, Mengyun Xu et al.· 0 citations
Offline reinforcement learning is limited by the coverage of the offline dataset, which makes it difficult for policies to generalize to unseen goals and behaviors. We introduce CoPL, a framework for counterfactual offline policy learning via large language models. Unlike prior relabeling approaches that modify only in...
Maosen Zeng, Yunan Liu· 2026 IEEE International Conf...· 0 citations
The Dual-Critic Constrained Deceptive Q-Learning (DCD-Q) method is proposed, a deployment-time trajectory protection framework that aims to reduce the information leaked by released trajectories while preserving acceptable task performance.
Guang-Yu Pan, Bo Hou, Yao Chen et al.· Journal of King Saud Univers...· 0 citations
Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to...
Xincong Hu, Lei Ou, Maosen Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.