Skip to content
Preprint

MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

Experiments demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.

Abstract

World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.

View source

Similar papers

Preprint Aug 2026

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Faster-WAM introduces a sparse future-conditioning framework that computes future representations once and selectively reuses them throughout action denoising, and proposes SparseMoT to replace ubiquitous layer-wise fusion with selective video-action interaction at a compact subset of network stages, and Interval KV-Fu...

Weiheng Zhao, Haoyi Jiang, Xin Shi et al. · 10 citations
Preprint Aug 2026

Foresight Without Seeing: Latent Futures for World Action Models

ForeWAM is proposed, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos, and demonstrates that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating futu...

Jiakai Huang, Zhongbo Wu, Siyu Xu et al. · 3 citations
Preprint Aug 2026

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives,...

Xiang Li, Yu-Peng Zheng, Song-En Gu et al. · 1 citation · ⚡1
Preprint Aug 2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Zongchuang Zhao, Xin Zhou, Tianyang Xu et al. · 4 citations
Preprint Aug 2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, whi...

Yi-Han Lin, Jia-Wei He, Shi-Feng Bao et al. · 10 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.