Skip to content
#generative ai Open access

PREreview of "Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23197406. ## Summary This paper studies Leni, a production enterprise AI business-analyst agent, and asks a question most vendor papers avoid: where does the system's measured reliability actually come from? Rather than reporting a single headline score, the author decomposes the uplift over the bare frontier base model across three public benchmarks chosen to stress different failure modes — SpreadsheetBench Verified (silent computation error), BullshitBench v2 (premise confabulation), and the GAIA validation split (cascade error over long tool chains). The full system improves by +11.0 percentage points on SpreadsheetBench (91.25% vs 80.25%, n=400, p<0.001), +7 to +10 points on BullshitBench (n=100), and roughly +15 points on GAIA validation (75.2% pass@1, n=165). The central finding is a decomposition: most of the uplift comes from scaffolding, routing, and specialist models rather than from the verification loop itself, whose isolated contribution is small (+1.5 points on SpreadsheetBench) but positionally decisive — it converts tasks at the top of the score distribution, the difference between mid-leaderboard and near-top. The author instruments the deterministic loop end to end, producing an empirical verifier confusion matrix (catch rate about 0.20, fix rate 0.75, no false-alarm regressions), and formalizes a compounding-reliability model with imperfect verification. Specialist-swap ablations suggest the observer matters as much as the loop: replacing the small post-trained verifier with the generating frontier model eliminates most rescues. A valid-premise control (100 expert-level questions, zero over-rejections) bounds the firewall's false-positive rate. ## Strengths 1. Honesty as a feature. The paper states which cross-system comparisons are statistically resolvable and which are not, and — remarkably — corrects its own company's previously published GAIA headline after re-grading every stored trajectory with the official scorer. That kind of self-correction in a vendor paper is rare and builds trust. 2. The decomposition is the right question for practitioners. Anyone who has shipped an agent knows the leaderboard number tells you nothing about what to build. Breaking the +11 points into scaffolding (+9.5) versus verification (+1.5) tells an engineering team where to spend effort. 3. The independent-observer finding rings true. My experience with self-critique in production matches this: a model checking its own work mostly agrees with itself. The specialist-swap result (rescues drop from 6 tasks to 2 when the frontier model verifies its own artifact) is the paper's most actionable insight. 4. A usable reliability model. Equation (1), p' = p(1−fb) + (1−p)cr, makes the verifier's error profile a first-class quantity. The condition under which a loop *reduces* reliability ((1−p)cr > pfb) is exactly the kind of guardrail a production deployment needs. ## Major comments 1. This is a vendor evaluating its own system. The limitations section says so, which I respect — but the decomposition rests on one production system and one team's scaffolding. Independent replication on a different stack is needed before treating "most uplift comes from scaffolding, not verification" as a general result. 2. No cost or latency accounting. Verification loops, specialist models, and multi-pass scaffolding multiply inference cost per task, yet the paper never reports dollars-per-task, token budgets, or latency. A production reader cannot decide whether the +1.5 points from the loop is worth the compute — the biggest gap in an otherwise production-minded paper. 3. GAIA validation split only (n=165). The author is candid about the corrected score, but the corrected pass@1 still sits on the validation split with a hidden-test-set submission listed as future work. The decomposition claims should be flagged as preliminary on GAIA. 4. Catch rate 0.20 deserves more attention. Four out of five true errors slip through the deterministic loop. The compounding model is elegant, but its sensitivity to the (c, r, f) estimates is never explored — if the catch rate drops to 0.10 on a different domain, does the loop still pay for itself? ## Minor comments 1. BullshitBench is a community benchmark (n=100) — the paper weights conclusions accordingly, which is appropriate. 2. The judge panel spans three providers but two configurations are Claude-based with a Claude judge; the paper flags judge bias, but a sensitivity check with the Claude judge excluded would strengthen this. ## Overall assessment Recommend with revisions. The decomposition, the honest uncertainty accounting, and the independent-observer result are genuinely valuable to anyone building production agents. I would ask for per-task cost/latency reporting and a clearer flag on the GAIA preliminary status before treating this as settled. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.