Skip to content

Reward Hacking Challenges Oversight of Autonomous Research Agents

Sep 2026 · 1 citation · 38 references
Computer Science

TL;DR

How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.

Abstract

Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack without instructions to do so, (2) how effective and detectable their methods are when hacking is allowed, and (3) how they adapt when an LLM review panel returns its decision and reasons. Across 17 language models and 38 tasks, the spontaneous reward-hacking rate is 30.5% on open-ended research-pipeline tasks and 2.9% on task-specific kernels. When hacking is allowed on tasks whose pass thresholds exceed our best compliant baselines, 505/677 attempts (74.6%) are confirmed reward hacks: they both clear the threshold and receive mechanism-verification panel confirmation of an evaluation exploit. An LLM panel reviewing only submitted code and reported scores misses 33/505 confirmed hacks (6.5%). Direct methods that achieve the highest scores are often easy to detect, while less direct methods evade more often. In a five-round loop, the number of model-task pairs with an evasion rises from 7 to 56. Among 79 pairs evaluated under two feedback conditions, cumulative evasion reaches 40.5% with detailed feedback and 20.3% with generic rejection. The detailed condition includes the review decision, reasons, and attempt history, so this comparison does not isolate the effect of explanations. These findings highlight the need for stronger defenses, including metrics kept outside the agent's control and independent recomputation on data chosen to expose likely exploits.

View source

Similar papers

Preprint Aug 2026

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

This work adapts HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduces Hack-Verifiable Terminal Bench (HVTB), to measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior.

Amit Roth, I. Bercovich, Yonathan Efroni · 0 citations
#artificial intelligence Preprint Aug 2026

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-t...

Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al. · 3 citations
#machine learning Preprint Sep 2026

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

This work finds that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max, and provides evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models.

Leon Bergen, Usha Bhalla, Andrew Lee et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

It is shown that selective control under partial auditing reduces accepted errors while increasing correctness, and experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing cor...

Christian Moya, E. Thornley, Guang Lin · 0 citations
Review Open access Aug 2026

A survey of reward hacking in agentic large language model systems

Large language models (LLMs) deployed as agentic systems capable of tool use, code execution, file manipulation, and multi-step planning inherit and amplify the classical reinforcement learning problem of reward hacking. This survey synthesizes how proxy-based alignment and evaluation failures manifest across modern LL...

Arun Morampudi, Ujval Irrinki, Rahul Grandhi et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.