A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a c...
Chu-Yao Fu, Xiao-Wei Chi, Yu-Han Rui et al.· 0 citations
RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning, is presented, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.
Rong-Yu Zhang, Rui-Zhi Fan, Yunfan Lou et al.· 0 citations
This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelit...
Yifan Ye, Yankai Fu, Ya-hui Lv et al.· arXiv.org· 4 citations
DynamicWAM introduces history-flow conditioning, encoding temporally aligned optical-flow frames alongside the current observation through a frozen pretrained video VAE to preserve spatial motion structure, while injecting kinematic descriptors of displacement, duration, velocity, and acceleration into the action exper...
Y. Lou, He-Wen Gao, Xiyu Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.