Skip to content
Preprint

An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM

Sep 2026 · 0 citations · 90 references
Computer Science

TL;DR

This work introduces PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch, which allows a single model to predict both how the robot should move and what the scene may look like afterward.

Abstract

Generative visual models offer a foundation for learning representations of physical dynamics, yet their extension to continuous control raises a fundamental question: do visual prediction and action generation require separate computational pathways? Existing approaches usually introduce trainable action heads or separate action experts to bridge low-dimensional states and high-dimensional visual representations. In this work, we explore whether the visual backbone's existing capacity can also support control when actions are expressed in a compatible representation. Thus, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch. This allows a single model to predict both how the robot should move and what the scene may look like afterward. Visual prediction and action generation become parts of the same generative process, without a dedicated action head or separate action expert. Experiments with subsampled training windows show gains over a matched dual-expert control, while benchmark evaluations reach 91.8% success rate on LIBERO-Plus and 96.12% on RoboTwin 2.0 in a full-data setting with additional augmented demonstrations. More broadly, the result suggests that capability need not be added where it can be inherited: the constraint on extending a generative backbone is the interface a new signal is written in, not the capacity to model it.

View source

Similar papers

Preprint Sep 2026

Rethinking Representations for World-Action Modeling

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e...

Hao-Yi Jiang, Liu Liu, Xin-Jiang Wang et al. · 1 citation
Preprint Sep 2026

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO succes...

Jian-Man Lin, S. Shailesh, Zhong-Yi Luo et al. · 1 citation
Preprint Oct 2026

Native Action-Prior Learning from Videos for World Action Models

World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual rep...

Zhao-Chong An, Fei Zhang, Meng-Lin Jia et al. · 0 citations
Preprint Sep 2026

From World Models to World Action Models: Rethinking Next-State Prediction

CF-WAM is proposed, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections.

Ting-Yu Yuan, Zi-Ming Ji, Biao-Liang Guan et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.