Skip to content

Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

This work examines how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error, and develops a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text.

Abstract

A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error. We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method. We also map representative defenses across substrates, identifying which mechanisms transfer directly and which offer only functional analogies. Persistent prompts receive particular attention: their contents are inspectable, but the behavior induced by a small textual change may be difficult to anticipate. The formal analysis, numerical illustration, and published evidence together provide a basis for comparing optimization methods and identifying the conditions under which their defenses transfer. The resulting framework connects optimization choices to verification requirements: reliable improvement depends on controlling accessible failure modes and preserving evidence of task quality independent of the score being optimized.

View source

Similar papers

Preprint Aug 2026

Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization

This work trains language models with GRPO on multiple-choice math problems where the correct answer is always option A, then evaluates on an unseen test set with unbiased answer positions to find reasoning-answer decoupling, which separates capability loss from a learned, transferable shortcut.

Suyash Maniyar, Armaan Sandhu, Abhishek Mishra · 0 citations
#artificial intelligence Preprint Sep 2026

Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End

How do you validate a cheap, deterministic proxy for an oracle that is expensive, rate-limited, and non-stationary? We present a protocol built on adversarial falsification gates (negative control, dose response, bounded amplification, duplication penalty, length neutrality) that define and select the proxy, fitted on...

Elisha Bajemon, André Rochet · 2 citations · ⚡1
#artificial intelligence Preprint Aug 2026

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

This work releases BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set, and releases an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-t...

Pradyumna Shyama Prasad, M. Anto, Leon Eshuijs et al. · 3 citations
#artificial intelligence Review Sep 2026

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

ProCTOR is described, a Teacher-Student loop in which a stateful orchestrator holds all tool access, stateless subagents diagnose failures and draft mutations they cannot apply, and a Teacher grades those mutations under five deterministic guardrails: hermetic sandboxes, capability-disjoint roles, acceptance checks tha...

Vansh Wahi · 2 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.