On-policy Verbal Distillation is introduced, a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations and suggests that retaining student-generated prefixes helps preserve explo...
Jing Xiong, Hui Shen, Shansan Gong et al.· arXiv.org· 8 citations
Experiments show that agentic review continuously improves PRs through a generate-review-revise loop, outperforms single-turn fixed-context review in both decision accuracy and resolve rate after revision, transfers beyond review to improve issue-resolution models, and enables effective and efficient test-time scaling.
Ruoyu Wang, Jierun Chen, Shaowei Wang et al.· arXiv.org· 4 citations
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward ha...
Yi-Ming Du, Yu-Xin Jiang, Tao Yuan et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.