A reconstruction-stable authorization (ReSA) is formulated over representation transitions, sink dependencies, consumed versions, and grant scope and implemented in APAS-Finder and freeze predictions before an independent sink oracle observes execution.
Abstract
Approval mechanisms have become a primary safeguard for consequential actions in LLM-agent software. Yet the action shown for approval is often not the object ultimately consumed: workflow reload, transcript projection, argument rebinding, and durable-state lookup may reconstruct it before execution. Existing fieldflow and check-coverage analyses can establish that expected fields were inspected, but not that the inspected object version reaches the sink or that no replacement intervenes between check and use. Consequently, a locally complete repair may still leave a residual authorization bypass after reconstruction. We address this problem through three designs. (1) We formulate reconstruction-stable authorization (ReSA) over representation transitions, sink dependencies, consumed versions, and grant scope. (2) We derive obligations that judge candidate repairs and expose residual sink suffixes. (3) We implement the analysis in APAS-Finder and freeze predictions before an independent sink oracle observes execution. Predictions agree with all 28 controlled outcomes. Four matched pairs require opposite judgments despite identical field-flow and check coverage; a CodeQL composite recovers all eight when supplied object flow, dominance, interference, and the same grant contract. Two analysts agree on all four model-validation dispositions and 19/20 sink dependencies. On released consumers, five repairs prevent 60/60 tested out-of-scope effects, while three mechanism contrasts expose the predicted residual effects. Repair sufficiency therefore depends on the reconstructed action consumed, not merely on an approval record or earlier checked representation.
EffectMatch is presented, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution and blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.
Hao-Ran Zhang, Heng-Tong Zhang, Zhi-Yu Liang et al.· 0 citations
Agents based on large language models (LLMs) can access heterogeneous devices through tools and APIs, but reliable execution must account for unmet effects, uncertain outcomes, and changing prerequisites. A command may be acknowledged without producing its intended effect, while missing feedback may obscure an action t...
Xue-Chun Li, Jia-Xin Liang, Jie Li et al.· 0 citations
Coding-agent approval interfaces bind a human decision to a command or tool call, while developer tools execute the transitive workflow that invocation activates. Package installation can run lifecycle hooks and write files; an MCP call can exercise network authority. We call the resulting record-coverage failure appro...
Jin-Qian Zhang, Hao-Jun Xia, Shu-Jiang Wu et al.· 0 citations
This work presents AID-Guard, a stateful authorization-to-effect closure protocol that revalidates the approved request and provider state at commit, retains one reservation under ambiguity, and permits release or one successor only after a terminal result or certified no effect with a delivery fence.
Large language model (LLM) agents increasingly execute long-horizon workflows through external tools, allowing untrusted outputs to influence subsequent actions and exceed user authorization. Existing defenses isolate injected content or constrain execution with predefined plans and static policies, but these approache...
Kai-Yuan Zhang, Yu-Ke Peng, Ke Jiang et al.· 1 citation· ⚡1
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.