Jul 2026
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
A staged length bonus is introduced that keeps reasoning length within a controlled range without simply encouraging brevity in multi-view spatial reasoning, and improves accuracy over vanilla GRPO while reducing average response length.
Xingjian Tao, Yiwei Wang, Yujun Cai et al.
· arXiv.org · 0 citations