Skip to content

When Does Execution Feedback Transfer? Minimal-Sufficient

· 0 citations · 9 references

TL;DR

This project studies that failure mode in a controlled Countdown arithmetic setting and asks whether verifier-generated execution feedback can supply a minimally sufficient correction signal that later transfers into a no-feedback policy, confirming the proposal’s central concern.

View source

Similar papers

Preprint Jul 2026

Verifier-Induced Support Reshaping in On-Policy Optimization

It is shown that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce, and effective rewardable support is defined as successful trajectories reachable within a fixed rollout budget.

Shaohang Wei, Z.Y. Su, Feifan Song et al. · 0 citations
Jul 2026

SCOPE-RL: Optimizing Reasoning Paths Before and After Success

SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO, indicating that reward-signal densification is complementary to policy-update-level RLVR advances.

Xiaojia Liu, Han Xu, Jianqiang Xia et al. · 0 citations
Jul 2026

RLPF: Reinforcement Learning from Performance Feedback for Code Generation

This work proposes RLPF, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward, and suggests that code agents can be trained not only to pass tests, but also to optimize the programs they write.

Huihao Jing, Haozhe Cui, Wenbin Hu et al. · 0 citations
Jul 2026

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

This work casts per-step rollout collection as a budget-constrained sequential allocation problem and introduces SARA (Sequential Adaptive Rollout Allocation), a two-threshold, SPRT-style rule that commits effective groups, abandons saturated ones after a short probe, and reallocates the freed budget to fresh prompts, without any extra prediction rollouts.

Pixel Nomand, Elena Voss, Marcus Hale et al. · 0 citations
Jul 2026

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

AdaKP is an online selector that re-chooses each problem's KP subset over the course of RL training, an entropy proxy that scores a KP by the reduction in next-token entropy it induces in a single inexpensive forward pass, with a provable bound on its truncation bias.

Zibin Meng, Zhenyu Zhao, Chunqiang Run · 0 citations
#reinforcement learning Preprint Aug 2026

AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification

This work proposes AutoVerifier, a residual-guided non-parametric optimization method that learns biases from recurring verifier errors and promotes them to code modules or prompt guidance only after replay validation detects no direct regressions, keeping accepted updates auditable, editable, and reusable.

Zelong Zhao, Zhihui Shi, Min-Qi Shi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.