This work finds that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max, and provides evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models.
Abstract
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations. Despite their simplicity, these vectors are both generalizable and interpretable, and we can use them to reliably detect reward hacking. We first evaluate reward hacking in commonly reported benchmarks like DeepSWE and SWE-bench, finding that models reward hack excessively in these environments; GLM 5.2 hacks in 57.2% of rollouts on DeepSWE and in 73% of rollouts on SWE-bench. Catching these requires monitors; LLM monitors are effective, but expensive detectors. We show that DoM vectors are similarly effective but virtually free, catching 3.1% more hacks in Kimi K3 and 7.9% fewer hacks in GLM 5.2 on DeepSWE at a monitor matched false positive rate. DoM vectors run on the chain-of-thought also predict reward hacks in the model's subsequent actions, meaning we can run them online and catch potential hacks before they occur. Finally, we analyze probe-hits that LLM monitors do not catch and discover other undesirable behaviors, as well as show transfer to finding hacks in non-SWE evaluations. Together, these results provide evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models
This work adapts HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduces Hack-Verifiable Terminal Bench (HVTB), to measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior.
Amit Roth, I. Bercovich, Yonathan Efroni· 0 citations
How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.
Yue Huang, Zhangchen Xu, Yu-Chen Ma et al.· 1 citation
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking d...
Shou-Li Wang, Yan-Feng Jia, Zhi-Hao Ou et al.· 0 citations
Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models. As an alte...
Oliver Daniels, Perusha Moodley, Benjamin M. Marlin et al.· 0 citations
This work examines how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error, and develops a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text.
This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-t...
Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al.· 3 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.