Skip to content

Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

EffectMatch is presented, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution and blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.

Abstract

Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can produce persistent effects beyond the requested change. Current safeguards can approve an action or record its aftermath, but without checking the persistent result before continuation, an unapproved outcome can be accepted as success and propagated to later steps. We present EffectMatch, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution. The comparison governs commit and dependent execution. In comparative evaluation on 206 public business tasks, EffectMatch preserved all clean executions and prevented all tested incorrect commits. Six 20-run ablations exposed the failure caused by each removed mechanism, while 80 task-topology cases preserved truthful handoffs and blocked invalid continuation. Together, these results show that EffectMatch blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.

View source

Similar papers

Preprint Sep 2026

ADF-EA: A Unified Execution Assurance System for Agent Device Foundation

Agents based on large language models (LLMs) can access heterogeneous devices through tools and APIs, but reliable execution must account for unmet effects, uncertain outcomes, and changing prerequisites. A command may be acknowledged without producing its intended effect, while missing feedback may obscure an action t...

Xue-Chun Li, Jia-Xin Liang, Jie Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ActGov: Governing LLM Agent Actions via Policy-Constrained Validation

Large language model (LLM) agents increasingly execute long-horizon workflows through external tools, allowing untrusted outputs to influence subsequent actions and exceed user authorization. Existing defenses isolate injected content or constrain execution with predefined plans and static policies, but these approache...

Kai-Yuan Zhang, Yu-Ke Peng, Ke Jiang et al. · 1 citation · ⚡1
#software testing Preprint Sep 2026

From Approval to Execution: Reconstruction-Aware Repair Analysis for LLM-Agent Software

A reconstruction-stable authorization (ReSA) is formulated over representation transitions, sink dependencies, consumed versions, and grant scope and implemented in APAS-Finder and freeze predictions before an independent sink oracle observes execution.

Jun-Chi Zhu, Zhen-Guang Liu, Shao-Jing Fan et al. · 0 citations
#artificial intelligence Preprint Oct 2026

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidenc...

Feng-Peng Li, Qi-Zhou Wang, Yu-Ke Hu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.