Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between vis...
Yu-Peng Zheng, Xiang Li, Song-En Gu et al.· 2 citations
WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos, achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa, and improves both policy performance and generalization in real-world manipulation tasks.
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives,...