This work introduces RoboFollow, a diagnostic benchmark with three principles, which exposes genuine instruction following as a critical, overlooked bottleneck in embodied agents' ability to follow instructions.
Chang Guo, Yu-Kun Xie, Bo-Han Tan et al.· 0 citations
Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard training. We propose an exploration-guided pr...
Yuan-Hao Yue, Qianli Ma, Cheng-Yu Wang et al.· 0 citations
Inspired by render-based compression, this work renders textual chains of thought into images, extract visual features, and construct a discrete latent vocabulary via clustering-based fine-tuning, and concludes that discrete latent tokens provide a controllable and interpretable basis for efficient latent reasoning.
Shuochen Chang, Qingyang Liu, Shaobo Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.