Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pro...
This work presents NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking that encodes full-environment obstacle and boundary geometry in a target-relative vertex representation, allowing the policy to retain route-defining context throughout the approach.
Zihan Wang, Baixiang Huang, Yang Guan et al.· 1 citation
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost o...
Shi-Qi Liu, Ze-Yu He, Le-Tian Tao et al.· 2 citations
This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.
B. Shuai, Min Hua, Le-Tian Tao et al.· Communications in Transporta...· 0 citations
A joint identifiability condition for controlled world models with Gaussian latent states with Gaussian latent states is presented, which consists of two coupled components: representation identifiability and transition identifiability, and it is proved that when this condition holds, minimizing the LeJEPA-style predic...
Xiangteng Zhang, Yang Guan, Bo Zhang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.