This project studies that failure mode in a controlled Countdown arithmetic setting and asks whether verifier-generated execution feedback can supply a minimally sufficient correction signal that later transfers into a no-feedback policy, confirming the proposal’s central concern.
It is shown that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce, and effective rewardable support is defined as successful trajectories reachable within a fixed rollout budget.
Shaohang Wei, Z.Y. Su, Feifan Song et al.· 0 citations
SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO, indicating that reward-signal densification is complementary to policy-update-level RLVR advances.
Xiaojia Liu, Han Xu, Jianqiang Xia et al.· arXiv.org· 0 citations
This work proposes RLPF, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward, and suggests that code agents can be trained not only to pass tests, but also to optimize the programs they write.
Huihao Jing, Haozhe Cui, Wenbin Hu et al.· arXiv.org· 0 citations
This work casts per-step rollout collection as a budget-constrained sequential allocation problem and introduces SARA (Sequential Adaptive Rollout Allocation), a two-threshold, SPRT-style rule that commits effective groups, abandons saturated ones after a short probe, and reallocates the freed budget to fresh prompts, without any extra prediction rollouts.
Pixel Nomand, Elena Voss, Marcus Hale et al.· arXiv.org· 0 citations
AdaKP is an online selector that re-chooses each problem's KP subset over the course of RL training, an entropy proxy that scores a KP by the reduction in next-token entropy it induces in a single inexpensive forward pass, with a provable bound on its truncation bias.
This work proposes AutoVerifier, a residual-guided non-parametric optimization method that learns biases from recurring verifier errors and promotes them to code modules or prompt guidance only after replay validation detects no direct regressions, keeping accepted updates auditable, editable, and reusable.
Zelong Zhao, Zhihui Shi, Min-Qi Shi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.