Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generate...
Shi-Long Zhou, Xiong-Hao Wu, Wen-Bo Li et al.· 0 citations
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-d...
Sen-Qiao Yang, Cheng-Yao Wang, Yuxin Chen et al.· 4 citations
Xpolicylab is released as shared infrastructure for reproducible policy comparison and standardized deployment across simulation and physical platforms for reproducible policy comparison and standardized deployment across simulation and physical platforms.
Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.
The results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.