Skip to content
Open access

Clean Attacks: Formalizing Semantically Valid Adversarial Behavior in Autonomous AI Agent Systems

2026 · International Journal of Scientific Research and Management · Vol 13, pp. 2623-2633 · 0 citations

TL;DR

A previously unstated class of adversarial input called a clean attack - an input that is syntactically correct, semantically consistent with the declared task context, consistent with all observable policy constraints and still has the goal of misguiding the agent away from the original operator goal - is identified and formalized in this paper.

Abstract

AI agents are being used more in high-pressure situations like managing email, running code, engaging with financial APIs, and supervising multi-agent pipelines. However, current taxonomy of adversarial attacks was mostly proposed for classifiers and generative models alone and fails to adequately describe the testbed of an agent with persistent state, multiple tools, and delegated power. A previously unstated class of adversarial input called a clean attack - an input that is syntactically correct, semantically consistent with the declared task context, consistent with all observable policy constraints, similar to legitimate operator instructions and still has the goal of misguiding the agent away from the original operator goal - is identified and formalized in this paper. These attacks go around the exposed dots of the “traditional” agent security architecture that only filters on the surface. The paper has three main contributions. One, it brings in a formal definition of the clean attack as a four-tuple of input, intent vector, policy envelope and behavioral outcome. Second, it suggests two operationalizable metrics: semantic validity score (SVS) and behavioral drift index (BDI) for systematically measuring the severity of clean attack. Third, the paper these metrics and taxonomy are validated, both by a purpose-built benchmark, AegisBench, and by 300 attack scenarios in three agent classes and (twelve) commercial agent pipelines. The experimental results show that clean attacks are a safety threat of a different category: while the conventional adversarial tasks are practically impervious to these attacks (4.2% success rate of the strongest agents), they achieve a mean attack success of 61.4%.

Read PDF

Similar papers

Preprint Aug 2026

Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures

It is argued that adversarial vulnerability stems from the absence of boundary verification, a security primitive that enforces explicit validation of data as it crosses inter-agent boundaries, including content, identity, execution intent, and state integrity.

Faisal Haque Bappy, Tahrim Hossain, T. S. Zaman et al. · 0 citations
Preprint Aug 2026

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

RedAgentBench is introduced, an executable framework for autonomous red-teaming and faithful measurement that shows that executable evaluation can improve safety measurement and identify actionable intervention points.

Zixing Chen, Xingyuan Liu, Jie Zhu et al. · 1 citation
Review Open access Sep 2026

Policy-Constrained Runtime Defense for Tool-Using AI Agents in Enterprise API Ecosystems

Tool-using AI agents can invoke internal APIs, retrieve documents, update records, and coordinate enterprise workflows. These capabilities create a runtime security problem: an agent may select an unauthorized tool, hallucinate an endpoint, follow malicious instructions embedded in retrieved context, rely on poisoned memory, retry unsafe operations, or submit a schema-valid but policy-violating payload. This paper presents a policy-constrained runtime enforcement framework that intercepts each proposed action before execution and classifies it as allow, deny, or escalate. We implement a deterministic trace-driven simulator with five service domains, six user roles, six threat classes, benign and adversarial tasks, and four defense configurations. The evaluation isolates enforcement effectiveness by replaying identical seeded action traces across all configurations. Across 8,000 controlled workflow executions, the framework reduces adversarial attack success from 100.0% for an unconstrained agent, 72.2% for prompt-only controls, and 18.5% for static gateway rules to 0.2%. It achieves a 99.9% overall safe-outcome rate, 100.0% benign safe completion under simulated reviewer approval, a 4.8% benign false-positive rate, and a 24.9 ms median enforcement latency. Ablation results show that registry validation, authorization, retry governance, and intent checking directly reduce attack success. Payload inspection addresses schema-valid semantic misuse, while context-integrity and escalation controls provide defense-in-depth and operational-governance benefits. The framework provides a structured basis for controlled evaluation of policy-constrained runtime enforcement across tool-using enterprise agents.

Swapneswar Ray · 0 citations
Open access Jul 2026

Invariant-Centered, Agent-Assisted Defense: A Security Architecture for Unbounded Attack Techniques and Unenumerable Attack Objectives

A defensive architecture for under unbounded attack techniques and unenumerable attack objectives, where the highest-value security investments are those that hold regardless of technique, and relocates residual reactivity to a single measurable point.

Ajayi Abisoye, N. Hussain, Abolaji Adebayo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.