Skip to content

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

It is shown that selective control under partial auditing reduces accepted errors while increasing correctness, and experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.

Abstract

In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.

View source

Similar papers

#machine learning Preprint Sep 2026

Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR

This work applies metamorphic testing to the verifier rather than the model, generating certified equivalent answer variants, that is, rewrites that preserve mathematical meaning by construction, so that any rejection is a provable false negative needing no human adjudication.

Esther Xin · 1 citation · ⚡1
#machine learning Preprint Sep 2026

CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL

During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking d...

Shou-Li Wang, Yan-Feng Jia, Zhi-Hao Ou et al. · 0 citations
#machine learning Review Sep 2026

Reward Hacking Challenges Oversight of Autonomous Research Agents

How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.

Yue Huang, Zhangchen Xu, Yu-Chen Ma et al. · 1 citation
#machine learning Preprint Aug 2026

Debate Training Reduces Reward Hacking in RLAIF

We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as training progresses, t...

Zachary Kenton, Lili Janzer, Rory Greig et al. · 1 citation
#machine learning Preprint Sep 2026

Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts

Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. Analysesbased on independent verifier errors may overlook dependence associatedwith shared answer formats. We investigate this dependence in24,998 groups of eight completions generated by Q...

Esther Xin · 1 citation
#machine learning Preprint Sep 2026

CircuitLens: Reasoning Circuits as Data Selection Signals for Reinforcement Learning with Verifiable Rewards

CRS is introduced, a selection signal derived from 46 reasoning-sensitive attention heads identified via contrastive ablation, computed in a single forward pass on the frozen base model without reward labels or rollouts, which runs against the intuitive hypothesis that stronger reasoning-circuit engagement produces bet...

Zhuo-Fan Chen, Zi-Qian Jiao, Yi-Kai Cui et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.