Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12...
Zhao-Chen Wang, Yu-Jun Cai, Huang-Bo Zou et al.· 0 citations
This work proposes SelectStream, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget.
Haonan Ge, Yi-Wei Wang, Hang Wu et al.· arXiv.org· 8 citations· ⚡1
A staged length bonus is introduced that keeps reasoning length within a controlled range without simply encouraging brevity in multi-view spatial reasoning, and improves accuracy over vanilla GRPO while reducing average response length.
Xingjian Tao, Yiwei Wang, Yujun Cai et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.