Skip to content

W2W: Language-Model-Based Trajectory Prediction with Reinforcement Learning

· 1 citation · 50 references

TL;DR

This work converts observed trajectories and interaction cues into parsable textual prompts, so that interaction semantics are expressed more explicitly in the model input and remains competitive with recent LM-based prediction methods and strong trajectory prediction base-lines on ADE/FDE.

View source

Similar papers

Preprint Aug 2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence.

Yihan Lin, Jiawei He, Shifeng Bao et al. · 2 citations
Preprint Aug 2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Zongchuang Zhao, Xin Zhou, Tianyang Xu et al. · 1 citation · ⚡1
Oct 2026

MaTF: Maneuver-Aware Temporal Fusion for Trajectory Prediction Under Arbitrary Observation Length

Trajectory prediction is essential for many robotic applications, yet most existing models rely on fixed-length observations and struggle with temporally irregular inputs. In real-world settings, prediction difficulty further increases when agents exhibit strong maneuverability, as their future motions depend on distinct short-term and long-term temporal cues. A Maneuver-aware Temporal Fusion framework is proposed to separate short-term dynamics from long-term intentions and fuse them through a motion-complexity-guided attention mechanism. The framework first extracts temporal features at different scales, and then adaptively balances them according to the maneuver patterns of each agent. To support incomplete or short observations, a self-distillation strategy is introduced to reconstruct missing motion segments, enabling consistent prediction without relying on explicit teacher-student models. Furthermore, a Mamba-Transformer hybrid backbone is employed to enhance computational efficiency and improve generalization under arbitrary observation lengths. Experiments on the ETH/UCY and SDD datasets show that MaTF consistently outperforms existing methods, particularly in scenarios with irregular or shortened observations.

Shuobo Wang, Wenyuan Qin, Yong-Zhao Hua et al. · 0 citations
#computer vision Preprint Aug 2026

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

Latent-OPD is proposed, which augments OPD with trajectory-level latent distillation and introduces a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers, establishing Latent-OPD as a highly effective approach to frame-efficient video reasoning.

Aoni Shen, Yongheng Zhang, Yinghui Li et al. · 1 citation
Preprint Aug 2026

Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach

This work proposes a novel test-time alignment approach that leverages trajectory-guided structured sampling for dynamic inference-time refinement, achieving better alignment with visual grounding and ensuring logical consistency, and demonstrates that this approach significantly improves accuracy without incurring prohibitive inference overhead.

Tianbao Jiang, Weicong Ni, Gerard de Melo et al. · 0 citations
Preprint Aug 2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy, which matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.

Jingkai Wang, Zihan Tang, Gu Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.