Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between vis...
Yu-Peng Zheng, Xiang Li, Song-En Gu et al.· 2 citations
VisCo is a training-efficient self-compression framework that reuses the pretrained VLM itself as an intrinsic compressor that compresses visual information using a small set of memory tokens and transfers hierarchical information from encoding to decoding.
Yupeng Zheng, Kai Zou, Bin Liu et al.· arXiv.org· 0 citations
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives,...
A latent memory pool is constructed that stores failure cases along with their structure scene representations and expert trajectory labels, and a dedicated Retrieve Model that decouples static road structure and dynamic agent interactions to enable structurally grounded retrieval is designed.
Zebin Xing, Yu-Peng Zheng, Qiang-Yu Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.