AR-WAM is presented, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt and a learnable operation token dictating the atomic skill to execute, predicting scene evolution within compact latent states while decoding actions.
Yi-Cheng Jiang, Ze-Sen Gan, Xiao-Bo Wang et al.· 0 citations
Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according to a gate. When the next gate is harvested from the previous accelerated forward, however, the tokens skipped at one step are also the ones...
Qi Luo, Shuaijun Liu, Hao Zhao et al.· 0 citations
JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, whi...
Yi-Han Lin, Jiawei He, Shi-Feng Bao et al.· 9 citations· ⚡1
LiLa-WAM is proposed, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU and the Visual Transition Token (VTT), a language-free task representation that encodes each task as a direction in visual feature space.
Fan Yang, Yu-Ting Su, Xiaobo Wang et al.· 9 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.