Skip to content

Fail Loudly: An Auditable Runtime for Agentic Data Analysis

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

RADAR is an auditable runtime that makes an agent's analytical choices inspectable and supports their revision through execution feedback and rejects the operation or provides diagnostic feedback, allowing the agent to revise its choices before errors propagate.

Abstract

Large language models (LLMs) have enabled data-science agents to automate multi-step analyses over heterogeneous files. However, incorrect choices regarding data sources, scope, or statistical definitions often lead to silent errors: computations execute successfully but produce plausible yet incorrect outputs that fail to answer the intended question. To mitigate this, we present RADAR, an auditable runtime that makes an agent's analytical choices inspectable and supports their revision through execution feedback. RADAR operates through three core mechanisms. First, an evidence-preserving exploration module retrieves task-relevant content while retaining source locations and observation coverage. Next, the runtime uses typed operators to record the agent's declared inputs, operation arguments, and resulting observations. Finally, runtime validation checks proposed operations against these observations. When a conflict is detected, the runtime rejects the operation or provides diagnostic feedback, allowing the agent to revise its choices before errors propagate. This design enables agents to fail loudly while leaving semantic interpretation to the LLM. On KramaBench, RADAR achieves overall scores of 0.723 with full source retrieval and 0.747 with gold sources supplied, corresponding to relative gains of 35.9% and 28.8% over the strongest baselines. Beyond KramaBench, RADAR achieves relative performance gains of 14.0% on DA-Code and 59.3% on DABStep, demonstrating its applicability across diverse agentic data-analysis workflows.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...

XiaoYu Xu, Zi Liang, Min-Xin Du et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

This work presents AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation and introduces AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks.

Wei-Da Liang, Shi Qiu, Zhun Wang et al. · 0 citations
Preprint Aug 2026

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

ABE-Ralph is introduced, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification.

Le-Zhi Yu, Xiao-Gang Xu, Yuhong Zhou et al. · 0 citations

Toward Self-Evolving Data Agents for Autonomous Data Analysis

Comparisons against stronger model and coding-agent competitors further indicate that both domain-specific agent runtime structure and foundation-model strength matter for autonomous data analysis.

Junhao Zhu, Lu Chen · 0 citations
Preprint Aug 2026

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, answer-consistent, and auditable.

Srimonti Dutta, Akshata Kishore Moharir · 0 citations
#artificial intelligence Preprint Sep 2026

Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation wit...

Param Raval, Rohit Shenoy, Archana Vaidheeswaran · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.