Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic fo...
Hui-Qiong Li, Zhi-Ting Mei, Anirudha Majumdar et al.· 1 citation
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while groundin...
Jia-Rui Yang, Jia-Jin Zhang, Bin Zhu et al.· 0 citations
Experiments demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.
VICAL is introduced, a consistency-driven framework that improves long-tailed recognition not by enforcing expert diversity, but by reducing prediction variance, and suggests that multi-expert models benefit more from variance reduction than diversity maximization.
Jianggang Zhu, Zheng Wang, Bin Zhu et al.· 0 citations
Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL.
Jia-Rui Yang, Bin Zhu, Jing-Jing Chen et al.· 1 citation
DECODE is proposed, a decoupled continual detection framework that jointly mitigates representation- and decision-level forgetting and introduces Subspace Diversity Regularization to preserve diverse forensic representations and Closed-Form Decision Alignment to recalibrate the shared classification head after each ada...
Zihao Cai, Xing-Hang Li, Ruiyan Yang et al.· arXiv.org· 0 citations
It is found that current models often generate visually coherent videos, but struggle with constraint reasoning, counterfactual grounding, physical interaction, and unsafe-instruction suppression, and results show that visual quality and surface-level instruction following are insufficient for trustworthy robotic video...
Huiqiong Li, Jia-Yu Wang, Zhiting Mei et al.· arXiv.org· 4 citations
Vorch-IR is presented, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model, and an automatic data construction pipeline that synthesizes paired supervision for all four editing settings is developed.
Yaowei Wang, Xiaoyu Chen, Xin Ma et al.· 0 citations
Saliency-guided Purification and Adaptive Redistribution (SPAR), a training-free, plug-and-play intervention that mitigates this generalized textual bias exerted over visual features that extends beyond isolated sink tokens.
Peng-Kun Jiao, Bin Zhu, Jingjing Chen et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.