This design-science study presents a protocol-agnostic architecture built around a canonical envelope, protected dependency and result hashes, Ed25519 or HMAC authentication, live-request rebinding, a seven-axis generated-claim gate, execution-time dependency revalidation, and eleven semantic invariants that support protected-dependency change detection, bounded outcome derivation, stale-decision prevention, surface-bound status consistency, refusal propagation, and verified-state identity.
Abstract
Agentic commerce extends agentic shopping into software agents that interpret policy, prepare checkout, generate transaction-facing language, and act under delegated payment authority. Protocols standardize external exchanges, but merchants still need one authoritative representation of commercial eligibility, actor authority, checkout validity, payment dispatch, generated claims, and evidence. This design-science study presents a protocol-agnostic architecture built around a canonical envelope, protected dependency and result hashes, Ed25519 or HMAC authentication, live-request rebinding, a seven-axis generated-claim gate, execution-time dependency revalidation, and eleven semantic invariants. Evaluation used an open-source JavaScript implementation, eight deterministic ecommerce scenarios, and five controlled ablations. Seven initially valid actions were permitted. After protected state changed, none could proceed without a fresh decision; a hostile-accessor case also remained blocked. Action status was consistent across configured surface-bound envelopes, and each scenario contained the three protected hashes and its targeted dependency reference. Each ablation produced the predicted unsafe regression when one safeguard was bypassed, while the protected path contained the same failure. The hostile accessor was read once, and the suite passed 66/66 tests, schema validation, and committed examples. Results support protected-dependency change detection, bounded outcome derivation, stale-decision prevention, surface-bound status consistency, refusal propagation, and verified-state identity under synthetic fixtures, but do not establish rule completeness, production security, performance, legal compliance, live interoperability, population error rates, or independent replication.
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.
Shichen Fan, Mingdai Yang, Duo Wang et al.· 0 citations
Agent payment protocols are emerging as a key transaction layer for autonomous commerce, enabling AI agents to purchase goods and services and execute payments on users'behalf. Unlike conventional payment flows, they distribute user intent, delegated authority, credential use, settlement, and fulfillment across multiple actors and stages, creating security dependencies that no single message or participant can enforce. Yet these guarantees remain largely implicit across evolving specifications, schemas, and reference implementations, with little systematic formal analysis. We formalize four representative agent payment protocols: x402, MPP, ACP, and AP2 in Tamarin. Using a common abstraction of the agent payment lifecycle, we construct source-grounded models that capture each protocol's roles, state, trust assumptions, and lifecycle transitions. Rather than assuming a complete property taxonomy, we use source-backed verification questions and counterexample traces to expose missing bindings, state constraints, and cross-stage correspondences, consolidating them into 18 shared security principles. Across 86 verification cases, our analysis reproduces 46 known or calibration cases and identifies 40 previously undocumented formal-consistency findings. For each retained violation, we isolate the missing protocol relation, construct a minimally strengthened reference model, and reverify the intended property. We further evaluate the new x402 findings across three implementations and validate ten representative findings through implementation PoCs, SDK/schema-level witnesses, and source-aligned executable traces spanning five security principles. Our results show that delegated authorization must remain consistent with its resulting economic and service effects across actors, states, and protocol stages.
Ke Jiang, Mo-Han Yu, Yuan-Yi-Chun-Min-Chieh Chang et al.· 0 citations
Payment networks and model providers deployed agent-authorization infrastructure at speed during 2025 and 2026: signed mandates, agent-bound tokens, and machine-payable settlement rails, each promising that an autonomous agent transacts only within authority its principal granted. This paper asks a prior question to whether agents obey such authority: whether the deployed protocols can express it at all. We define an authorization envelope of eight fields drawn from the delegated-authority literature and from the control primitives of existing payment rails, comprising a per-transaction ceiling, a cumulative ceiling, a merchant set, a category set, required product attributes, a validity window, a substitution policy and an amount-valued confirmation threshold. We then code eight deployed agent-payment protocols against these fields using an auditable document-analysis protocol, classifying each field as expressible, advisory or absent according to whether a typed schema field exists and whether any identified party validates it. Three fields are unsupported almost everywhere: substitution policy, general product attributes, and the confirmation threshold. Cumulative ceilings are enforceable only where some party accumulates state across transactions, which five of the ten schemes examined do and the remainder do not. Most consequentially, virtual-card controls already enforce cumulative caps and merchant-category scope, and open-banking variable recurring payments enforce cumulative caps, that the new agent protocols omit, so agent authorization is in specific respects a regression against rails that preceded it. We release the coding protocol and evidence table, and retain version-pinned specification snapshots for audit.
Ian Staley· Journal of Artificial Intell...· 0 citations
A policy algebra is proposed that defines the reliability envelope within which agent capability may be exercised and provides researchers and practitioners with formal correctness conditions, executable decision semantics, and trace evidence for building agents that are not only capable, but reliably capable.
Bhaskar Tripathi, Anurag Kumar, R. Kumar et al.· 0 citations
Agentic artificial intelligence expands the enterprise security boundary because autonomous agents can plan
tasks, retain memory, invoke tools, call APIs, and initiate business actions. Authentication at session start is therefore
insufficient when later actions may be influenced by untrusted content, poisoned memory, compromised tools, or excessive
delegated privilege. This paper proposes the Zero Trust Agentic AI Security Framework (ZT-AASF), a vendor-neutral
architecture that applies continuous verification to consequential agent actions. The framework separates six control
planes: identity and delegation, context and data trust, policy and risk decision, tool and action enforcement, runtime
observability, and containment and recovery. A contextual authorization model evaluates delegated scope, source
provenance, data sensitivity, tool risk, behavioral deviation, and action impact before execution. A design-level evaluation
against ten OWASP agentic risk classes produces 27 of 30 control-coverage points for ZT-AASF versus 5 of 30 for a
conventional integration baseline. These values represent architectural coverage, not measured attack-prevention rates.
The results indicate that moving enforcement from the session boundary to the action boundary can reduce implicit trust,
constrain privilege propagation, and improve auditability while preserving useful autonomy.
Sachin Suryawanshi· International Journal of Inn...· 0 citations
SAGE-Fin is presented, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, the object of runtime control, and its results establish executable conformance, not independent safety accuracy.
Rui Tang, Qiang Liu, Yichi Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.