This work adapts HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduces Hack-Verifiable Terminal Bench (HVTB), to measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior.
Abstract
As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack-verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduce Hack-Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward-hacking strategies, but also'unknown unknown'exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack-verifiable-environments/hvtb
How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.
Yue Huang, Zhangchen Xu, Yu-Chen Ma et al.· 1 citation
This work finds that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max, and provides evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models.
Leon Bergen, Usha Bhalla, Andrew Lee et al.· 0 citations
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking d...
Shou-Li Wang, Yan-Feng Jia, Zhi-Hao Ou et al.· 0 citations
This work evaluates escalation channels, structured reporting tools available to the agent at the point of conflict, as a decision-environment intervention that both reduces reward hacking and surfaces the infrastructure defects that trigger it.
This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-t...
Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al.· 3 citations
This work examines how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error, and develops a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text.
Vansh Wahi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.