Skip to content

When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

An effect-history model that separates events in the external world from the runtime's observations of them, and a catalog of eight recurring external-effect anomalies that motivate reusable transactional contracts at the tool boundary.

Abstract

AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the resulting external state may be inconsistent with the workflow's intended resolution: required effects may be missing or duplicated, aborted effects may survive, and committed effects may depend on provisional state that is later withdrawn. Advanced transaction models address related failures, but assume that lower-level operations expose the semantics they depend on: whether an effect occurred, whether it can be compensated, staged, or safely reordered. Shared agent-tool interfaces usually do not. We contribute an effect-history model that separates events in the external world from the runtime's observations of them, and a catalog of eight recurring external-effect anomalies. From the catalog we derive the boundary capabilities required to exclude each anomaly in general, and four points where black-box tool invocation alone cannot provide a general guarantee. We then ask how much of this is expressible in a widely used shared tool interface, measuring the use of the standard annotation vocabulary across 98,291 tools exposed by registered Model Context Protocol (MCP) servers. The fields are widely emitted but provide only coarse call-level hints, and none of the required capabilities is fully expressible. These results motivate reusable transactional contracts at the tool boundary.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows

EffectMatch is presented, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution and blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.

Hao-Ran Zhang, Heng-Tong Zhang, Zhi-Yu Liang et al. · 0 citations
Preprint Aug 2026

Metis: Typed Runtime Mediation for Tool-Using Software Agents

Metis, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects is presented, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects.

Jun'an Yu · 0 citations
#artificial intelligence Review Sep 2026

When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents

Formation-consistent dispatch (FCD), which connects implementation analysis to execution authority, is presented, which produces provenance-bound over-approximations of declared in-scope effects from official source.

Geonwoo Kim, Brent ByungHoon Kang Korea Advanced Institute of Science, Technology · 0 citations
#artificial intelligence Preprint Sep 2026

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents

LIMBO, a deterministic sandbox of six services with realistic contracts and twelve fault modes injected at the service boundary and twelve fault modes injected at the service boundary, is introduced, proving that no verification-only policy is exactly-once under late commits without a bound on in-flight time.

Jia-Peng Li · 1 citation
#artificial intelligence Preprint Oct 2026

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidenc...

Feng-Peng Li, Qi-Zhou Wang, Yu-Ke Hu et al. · 0 citations
Preprint Aug 2026

When"Must"Becomes"Maybe": Constraint Weakening in LLM Agent Workflows

This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.

Yi-Heng Sun, Hui-Fei Wang, Yan-Cheng Zhu et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.