Skip to content

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Sep 2026 · 1 citation
Computer Science

TL;DR

Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate; and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

Abstract

Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $\tau^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...

XiaoYu Xu, Zi Liang, Min-Xin Du et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the...

Hai-Qing Li, Xin-Yu Ma, Yin-Hao Wu et al. · 0 citations
Review Sep 2026

Evaluating Agent Skills for Version-Specific Plugin Migration: A Retrospective Study

The study contributes a traceable evaluation that connects aggregate reward to contract-level evidence and judge sensitivity, together with concrete review checks for migration advice, together with concrete review checks for migration advice.

Bei-Ming Liu, Hai-Hao Li, Min-Jie Chen et al. · 0 citations
Preprint Aug 2026

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

Evidence-Carrying Termination (ECT): an agent may return COMPLETE only when a typed certificate binds every required answer claim to valid, in-scope trace evidence and a deterministic replay reconstructs the claimed value.

Jason Liu · 3 citations
#artificial intelligence Preprint Sep 2026

Dude, Where's My State? Execution Information Requirements for Stateful Agents

Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with kno...

N. Mehrotra, Ashish Tiwari, Priyanshu Gupta et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.