An effect-history model that separates events in the external world from the runtime's observations of them, and a catalog of eight recurring external-effect anomalies that motivate reusable transactional contracts at the tool boundary.
Abstract
AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the resulting external state may be inconsistent with the workflow's intended resolution: required effects may be missing or duplicated, aborted effects may survive, and committed effects may depend on provisional state that is later withdrawn. Advanced transaction models address related failures, but assume that lower-level operations expose the semantics they depend on: whether an effect occurred, whether it can be compensated, staged, or safely reordered. Shared agent-tool interfaces usually do not. We contribute an effect-history model that separates events in the external world from the runtime's observations of them, and a catalog of eight recurring external-effect anomalies. From the catalog we derive the boundary capabilities required to exclude each anomaly in general, and four points where black-box tool invocation alone cannot provide a general guarantee. We then ask how much of this is expressible in a widely used shared tool interface, measuring the use of the standard annotation vocabulary across 98,291 tools exposed by registered Model Context Protocol (MCP) servers. The fields are widely emitted but provide only coarse call-level hints, and none of the required capabilities is fully expressible. These results motivate reusable transactional contracts at the tool boundary.
EffectMatch is presented, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution and blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.
Hao-Ran Zhang, Heng-Tong Zhang, Zhi-Yu Liang et al.· 0 citations
Metis, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects is presented, a multi-provider runtime that converts provider streams into typed events before admitted calls reach external effects.
Formation-consistent dispatch (FCD), which connects implementation analysis to execution authority, is presented, which produces provenance-bound over-approximations of declared in-scope effects from official source.
Geonwoo Kim, Brent ByungHoon Kang Korea Advanced Institute of Science, Technology· 0 citations
LIMBO, a deterministic sandbox of six services with realistic contracts and twelve fault modes injected at the service boundary and twelve fault modes injected at the service boundary, is introduced, proving that no verification-only policy is exactly-once under late commits without a bound on in-flight time.
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidenc...
Feng-Peng Li, Qi-Zhou Wang, Yu-Ke Hu et al.· 0 citations
This work identifies a state-transmission failure between information extraction and action in large language model agents, and shows how handoff transformations can retain state content while weakening its constraints on downstream action.
Yi-Heng Sun, Hui-Fei Wang, Yan-Cheng Zhu et al.· 2 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.