Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embo...
Hao Wang, Jia-Jun Wen, Jing-Zhi Liu et al.· 0 citations
Retrieve in Time, Correct in Frequency (RTCF), a training-free test-time correction framework that improves frozen VLA performance with low model-side overhead and requires no parameter updates, repeated VLA inference, or additional GPU resources.
Yuze Fan, Yue Cao, Peng-Jie Gao et al.· 0 citations
Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challeng...
Yuanruyi, Yue Cao, Haojia Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.