Skip to content

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

Sep 2026 · 0 citations · 89 references
Computer Science

TL;DR

This work finds that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max, and provides evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models.

Abstract

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations. Despite their simplicity, these vectors are both generalizable and interpretable, and we can use them to reliably detect reward hacking. We first evaluate reward hacking in commonly reported benchmarks like DeepSWE and SWE-bench, finding that models reward hack excessively in these environments; GLM 5.2 hacks in 57.2% of rollouts on DeepSWE and in 73% of rollouts on SWE-bench. Catching these requires monitors; LLM monitors are effective, but expensive detectors. We show that DoM vectors are similarly effective but virtually free, catching 3.1% more hacks in Kimi K3 and 7.9% fewer hacks in GLM 5.2 on DeepSWE at a monitor matched false positive rate. DoM vectors run on the chain-of-thought also predict reward hacks in the model's subsequent actions, meaning we can run them online and catch potential hacks before they occur. Finally, we analyze probe-hits that LLM monitors do not catch and discover other undesirable behaviors, as well as show transfer to finding hacks in non-SWE evaluations. Together, these results provide evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models

View source

Similar papers

Preprint Aug 2026

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

This work adapts HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduces Hack-Verifiable Terminal Bench (HVTB), to measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior.

Amit Roth, I. Bercovich, Yonathan Efroni · 0 citations
#machine learning Review Sep 2026

Reward Hacking Challenges Oversight of Autonomous Research Agents

How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.

Yue Huang, Zhangchen Xu, Yu-Chen Ma et al. · 1 citation
#machine learning Preprint Sep 2026

CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL

During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking d...

Shou-Li Wang, Yan-Feng Jia, Zhi-Hao Ou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Inducing Emergent Misalignment from Reward Hacks with Iterative DPO

Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models. As an alte...

Oliver Daniels, Perusha Moodley, Benjamin M. Marlin et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-t...

Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al. · 3 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.