HANSARD is proposed, a reference architecture treating accountability as a life-cycle property, i.e., spreading an act across redundant agents until none is a but-for cause, making laundering visible.
Abstract
Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations. Already, the first largely AI-orchestrated intrusion campaigns have been reported. Yet, when such a system causes harm, no method can robustly establish what happened, what caused it, or who is accountable. This is because provenance forensics works at the wrong abstraction, formal causality assumes the causal model, and agent auditing trusts self-recording. The target failure mode is, thus, attribution laundering, i.e., spreading an act across redundant agents until none is a but-for cause. Worse, the record is produced by the suspects, which comprises the assumption adopted throughout this work. Agents may therefore anticipate the investigation and the part of logging infrastructure may itself collude. In this paper, HANSARD is proposed, a reference architecture treating accountability as a life-cycle property. First, a readiness profile sealed before operation bounds what later findings may claim. Second, capturing at five choke points beyond the agents'reach makes omissions detectable, not only tampering. Third, a typed PROV-DM-aligned causal graph accrues as the system runs, and three indicators read it live to gate oversight without adjudicating. Fourth, post-incident replay yields contingent effects under the modified Halpern-Pearl definition, together with a compensation-set size. Finally, a synergy residual measures harm due to the combination rather than to individuals, making laundering visible. Cause, responsibility and accountability are then reported separately, each capped by an evidentiary tier, while a future research agenda is also provided.
A pipeline promoting an AI system publishes records claiming the thing evaluated is the thing deployed and that the evidence licensed the transition, and measures whether those records can express that claim and whether it holds where declared.
This work argues that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns.
Kai Wang, Zeming Wei, Biaojie Zeng et al.· 0 citations
Autonomous AI agents now hold execution authority over high-consequence enterprise actions in finance, healthcare, and infrastructure, where catastrophic failures are rare yet dominate systemic risk and probabilistic content filtering does not constitute a reference monitor over what is executed. This paper introduces L-DREA, a deterministic runtime-enforcement architecture that generalizes Anderson’s 1972 reference-monitor primitive from mediation of data access to mediation of externally effective action. L-DREA separates capability generation from execution authority, binds every candidate action to an epoch-keyed Permit-to-Act token, and interlocks externalization through a commit-before-actuate substrate. Five structural properties generalize Anderson’s primitive — complete mediation, tamper-resistance, verifiability, non-compensatory aggregation, and epistemic bounding — and six runtime invariants are established analytically, the first additionally mechanized in TLA+ with a released TLC log. A software (Tier-S) reference implementation is evaluated on two disjoint evidence tracks: a seeded synthetic corpus of 1,217,906 runtime proposals (360,000 adversarial), and the public ULB credit-card dataset (284,807 transactions) as a golden-oracle authorization trace, plus a blind committed-before-label-reveal protocol on three public datasets. Across both tracks, a 120,000-attempt full-knowledge adaptive attacker, 2,394 injected runtime attacks, live revocation and watchdog suites, and an executed offline AgentDojo run, zero unauthorized externalizations were observed; exact one-sided Clopper–Pearson and Wilson upper bounds accompany every zero-event claim, with replay determinism of 100.0000% over 1,217,906 cycles and 24,912 Ed25519-signed permit tokens runtime-verified at 100% integrity. All results are bounded to the documented threat surface and released as a seeded, one-command reproducible artifact; hardware substrates and HSM key custody are specified, not claimed.
Multi-agent systems (MAS) comprise autonomous software agents that collaborate to perform complex tasks in critical cyber-physical domains, including multi-robot coordination and the Industrial Internet of Things (IIoT). In such distributed environments, a compromised agent may execute modified software while appearing trustworthy, causing other agents to act on false information and corrupting the mission. Agents must therefore establish and maintain mutual trust throughout operation. Remote attestation (RA) is a well-established technique for this purpose, enabling a remote verifier to assess the integrity of a potentially compromised prover device. However, conventional RA approaches face significant limitations in MAS: integrity guarantees are restricted to boot or application-load time, designs rely on centralized trusted verifiers or security hardware, and attestation records lack transparency and auditability. To address these limitations, this paper presents D-MUTRA, a blockchain-based framework that introduces a mutual RA protocol in which agents measure their runtime integrity while verifying that of their peers, acting as both prover and verifier. The framework operates entirely in software and relies on two components: a Security-as-a-Service that instruments agents with lightweight measurement and verification capabilities, and a smart contract that coordinates the attestation protocol in a decentralized and transparent manner. We implement a proof-of-concept on a private Ethereum blockchain using Hyperledger Besu and evaluate it in a swarm robotics scenario built with Robot Operating System (ROS) and the Gazebo simulator. Results show that D-MUTRA enables agents to continuously attest one another, detects malicious software modifications, and scales to large deployments with negligible overhead on protected applications.
Adam Zahir, V. Lefebvre, M. Angoustures et al.· 0 citations
Canonical Action Verification and Attestation (CAVA), a runtime-semantics layer for converting heterogeneous agent activity into canonical runtime action objects, is presented.
Observability platforms have historically been read-only consumers of telemetry: they collected, correlated, and displayed. Recent reference architectures for generative and agentic root cause analysis (RCA) break that assumption. When an initial diagnosis falls below a confidence threshold, the system invokes an agent that autonomously queries monitoring tools, reads source repositories, and executes policy-driven actions until sufficient evidence is assembled. The observer has become an actor. This paper argues that the security consequences of that shift have not been systematically analyzed, and supplies the missing analysis. Taking the four-layer GenAI-driven observability architecture of [1] as a representative system under analysis, we construct an explicit threat model: three trust zones, three trust boundaries, four adversary classes, and nineteen threats enumerated with STRIDE across seven architectural elements. We then trace three end-to-end attack chains that each cross all three boundaries, the most consequential being telemetry-borne instruction injection, in which an adversary who can influence a single application log line reaches the production control plane without ever authenticating to the observability platform. Two structural properties make agentic observability distinctively exposed: telemetry is attacker-influenced input that has never been treated as such, and the diagnostic agent necessarily holds broad standing read privilege across the estate. We propose twelve design-level controls organized into four families — provenance, corpus integrity, least authority, and accountability — and identify the residual risks that no current control fully addresses. This work is analytical in scope. It contributes no implementation and reports no experiments; its purpose is to establish the threat vocabulary and design constraints that empirical work in this area will need.
Shiva Raju· International Journal For Mu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.