Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...
Jia-Yi Chen, Wen-Xuan Song, Jing-Bo Wang et al.· 0 citations
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...
Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al.· 0 citations
4D-WAM is proposed, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment, enabling WAMs to learn trajectory-level spatiotemporal representations.
Lishan Yang, Wen-Xuan Song, Xi Wang et al.· 5 citations· ⚡1
PIVOTS is introduced, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs'ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research and examines how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS d...
Shuxiang Zhang, Yiting Yin, Wen-Xuan Song et al.· arXiv.org· 0 citations
Xpolicylab is released as shared infrastructure for reproducible policy comparison and standardized deployment across simulation and physical platforms for reproducible policy comparison and standardized deployment across simulation and physical platforms.
MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization.
Ze-Hua Fan, Jun-Jie He, Wen-Xuan Song et al.· 3 citations· ⚡1
DreamTrajectory is presented, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation of existing Vision-Language-Action policies, and jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert.
Zheng Yang, Wen-Jie Zhang, Xiang-Yu Chen et al.· 1 citation
Xiao-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency and across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods.
Xiaomin Guo, Piao-Piao Jin, Jason Li et al.· arXiv.org· 16 citations· ⚡2
MoWorld is the first real-time interactive World Model built on the Neural Processing Unit (NPU) and can achieves up to 50 FPS in such the devices, enabling practical and efficient deployment at scale.
Team Moxin, Deyi Ji, Tianrun Chen et al.· arXiv.org· 1 citation
The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.
Hao-Dong Yan, Jun-Feng Li, Jun-Jie He et al.· 2 citations
PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.