Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where...
Ding-Sheng Liu, Yang-Zheng Wu, Mahboubeh Asadi et al.· 0 citations
Faster-WAM, an instantiation of DoT for WAMs, which docks a single-layer action head onto a 30-layer video backbone, achieves competitive performance on LIBERO and RoboTwin 2.0 while demonstrating strong out-of-distribution generalization on LIBERO-Plus.
Liheng Ma, Rui Yang, Zhan-Guang Zhang et al.· 4 citations
KAM-WM is presented, a framework that extracts a coarse directional interaction cue from a frozen latent video world model without rollout or world-model fine-tuning and indicates that a frozen video model can provide a useful first-order visual prior for control without the test-time cost of future rollout.
RoboHarness is proposed, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills as well as a general framework compatible with a broader range of robot policies, such as navigation policies, model predictive controllers, and world-action models.
Jinbang Huang, Yuan Hu, Zhiyuan Li et al.· arXiv.org· 3 citations· ⚡1
Future-State-Conditioned VLN (FSC-VLN), a deployable model that augments a causal policy with a future-query token and uses training-only future-state supervision to distill information from future observations into the policy state, is proposed.
Lingfeng Zhang, Zhanguang Zhang, Liheng Ma et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.