Skip to content
Preprint

DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

This work introduces DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction, each token is a single vector encoding changes between consecutive dense DINO feature maps.

Abstract

World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between them are what an action policy needs to anticipate. We introduce DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction. Each token is a single vector encoding changes between consecutive dense DINO feature maps. DeltaWAM builds on DeltaWorld, a latent world model pretrained on large-scale videos, to autoregressively predict one delta token per future frame. A flow-matching action expert then conditions on the predicted transitions and current DINO features, which serve as spatial anchors, to generate action chunks. Trained for 256 GPU hours on two H100 GPUs, DeltaWAM has 0.725B parameters and achieves 92.8% average success on LIBERO. It also shows robust generalization under procedural perturbations on LIBERO-Pro. Inference takes 142.1 ms per action chunk with 3.86 GB peak memory.

View source

Similar papers

Preprint Sep 2026

DeltaWAM: Delta World Action Models for Bimanual Manipulation

This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas,...

Han Yan, Zi-Shang Xiang, Hao-Kai Jiang et al. · 0 citations
Preprint Aug 2026

Foresight Without Seeing: Latent Futures for World Action Models

ForeWAM is proposed, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos, and demonstrates that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating futu...

Jiakai Huang, Zhongbo Wu, Siyu Xu et al. · 3 citations
Preprint Aug 2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.

Hao-Dong Yan, Jun-Feng Li, Jun-Jie He et al. · 2 citations
Preprint Sep 2026

CoRe-WAM: Correspondence-Aligned Temporal Residuals for World Action Models

Comparing current and past observations helps robots understand scene changes and select subsequent actions during manipulation. However, comparing visual features at the same image location can mix different scene content when objects or the camera move. We introduce CoRe-WAM, a world-action model that incorporates co...

Bing-Heng Zhou, Jia-Long Liu, Jia-Nan Wang et al. · 0 citations
Preprint Aug 2026

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

LD4WAM is presented, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated fut...

Zhen Shen, Jia-Qi Liang, Jasper Lu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.