Skip to content
Preprint

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

This work adapts HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduces Hack-Verifiable Terminal Bench (HVTB), to measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior.

Abstract

As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes an increasingly important failure mode. Measuring reward hacking is itself challenging, as detection typically relies on human inspection or LLM judges, both of which can be unreliable. The hack-verifiable environments (HVE) methodology addresses this challenge by embedding detectable hacks into tasks, allowing reward hacks to be identified automatically and reliably. In this work, we adapt HVE to Terminal Bench, a leading benchmark of real-world terminal and coding tasks, and introduce Hack-Verifiable Terminal Bench (HVTB). Using HVTB, we measure reward-hacking rates across frontier models and study whether prompts with varying amounts of information on the hack can mitigate this behavior. This lets us test whether prompting can prevent not only known reward-hacking strategies, but also'unknown unknown'exploits that the prompt does not anticipate. We release all environments and agent traces at https://majoroth.github.io/hack-verifiable-environments/hvtb

View source

Similar papers

#machine learning Review Sep 2026

Reward Hacking Challenges Oversight of Autonomous Research Agents

How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.

Yue Huang, Zhangchen Xu, Yu-Chen Ma et al. · 1 citation
#machine learning Preprint Sep 2026

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

This work finds that simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max, and provides evidence that simple, white-box methods can be used to scalably study and monitor reward hacking behaviors in frontier open source models.

Leon Bergen, Usha Bhalla, Andrew Lee et al. · 0 citations
#machine learning Preprint Sep 2026

CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL

During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking d...

Shou-Li Wang, Yan-Feng Jia, Zhi-Hao Ou et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-t...

Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.