Jul 2026
SCOPE-RL: Optimizing Reasoning Paths Before and After Success
SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO, indicating that reward-signal densification is complementary to policy-update-level RLVR advances.
Xiaojia Liu, Han Xu, Jianqiang Xia et al.
· arXiv.org · 0 citations