This work introduces DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction, each token is a single vector encoding changes between consecutive dense DINO feature maps.
Abstract
World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between them are what an action policy needs to anticipate. We introduce DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction. Each token is a single vector encoding changes between consecutive dense DINO feature maps. DeltaWAM builds on DeltaWorld, a latent world model pretrained on large-scale videos, to autoregressively predict one delta token per future frame. A flow-matching action expert then conditions on the predicted transitions and current DINO features, which serve as spatial anchors, to generate action chunks. Trained for 256 GPU hours on two H100 GPUs, DeltaWAM has 0.725B parameters and achieves 92.8% average success on LIBERO. It also shows robust generalization under procedural perturbations on LIBERO-Pro. Inference takes 142.1 ms per action chunk with 3.86 GB peak memory.
This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas,...
Han Yan, Zi-Shang Xiang, Hao-Kai Jiang et al.· 0 citations
ForeWAM is proposed, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos, and demonstrates that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating futu...
Jiakai Huang, Zhongbo Wu, Siyu Xu et al.· 3 citations
Causal Imprint is introduced, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference.
Xing-Yu Miao, Zi-Zun Li, Bao-Le Fang et al.· 0 citations
The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.
Hao-Dong Yan, Jun-Feng Li, Jun-Jie He et al.· 2 citations
Comparing current and past observations helps robots understand scene changes and select subsequent actions during manipulation. However, comparing visual features at the same image location can mix different scene content when objects or the camera move. We introduce CoRe-WAM, a world-action model that incorporates co...
Bing-Heng Zhou, Jia-Long Liu, Jia-Nan Wang et al.· 0 citations
LD4WAM is presented, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated fut...
Zhen Shen, Jia-Qi Liang, Jasper Lu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.