Skip to content

Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

The studies address different limits: what an execution comparison estimates, and what target recall captures are addressed: what an execution comparison estimates, and what target recall captures.

Abstract

What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 24-day upward phase mixes the execution rule with fresh model responses and portfolio feedback: the parsed decision paths agree in only $19.8\%$ of $450$ pairs. Replaying each stored response tape through both execution destinations gives a narrower result. Conditional on those responses, stressed execution changes total return by $-0.0170$ (95\% interval $[-0.0230,-0.0117]$), or $10.4\%$ of the idealized baseline, and ten seed clusters do not resolve the model ranking. Study~B corrects an incomplete answer key and replaces legacy tasks with matched zero-, one-, and two-defect tasks under an explicit multi-label prompt. The drop in target violation recall from one to two defects is positive in five of six combinations of auditor and source (median $0.267$), with three surviving Holm correction. Yet the auditor that includes both target labels most often has micro-precision $0.149$, emits findings on $98/100$ zero-defect tasks, and returns the exact dual-defect set in only $21/100$ cases. Target recall by itself therefore gives a poor account of audit quality on this construction. The studies address different limits: what an execution comparison estimates, and what target recall captures. Together, they show how fixed conditions and diagnostic controls bound the claims a score can support.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate; and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction...

Yu-Kun Zhang, Ke-Mu Xu, Yi-Shen Chen · 1 citation
Preprint Aug 2026

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, are introduced to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence to frame financial compliance evaluation as an audit of rule-grounded actions and eviden...

Yiyan Luo, Yihang Jiang, Qijun Xie et al. · 1 citation
#artificial intelligence Preprint Sep 2026

The Delegation Blind Spot: Auditing Product Decisions from Agent Choices

Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. Its foundations are established identification and decision theory...

Shivam Gupta · 0 citations
#machine learning Preprint Sep 2026

Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation

The fork ledger is introduced, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers, allowing triggers to be judged by the updates they select rather than by surprise detection alone.

An-Qi Li, Kaden Kim · 0 citations
Open access Sep 2026

Stateful Mediation and Selective Auditing for Financial Language-Model Actions: A Paired Controlled Simulator Study

Language-model agents can propose financial actions based on observations that become stale before execution. This study measures how five execution arrangements translate the same model proposals into post-state harm and correct completion in a controlled synthetic financial workflow. A protocol was internally frozen...

A. Ojha · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.