Skip to content

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

This work presents AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation and introduces AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks.

Abstract

AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.

View source

Similar papers

Review Open access Sep 2026

Policy-Constrained Runtime Defense for Tool-Using AI Agents in Enterprise API Ecosystems

Tool-using AI agents can invoke internal APIs, retrieve documents, update records, and coordinate enterprise workflows. These capabilities create a runtime security problem: an agent may select an unauthorized tool, hallucinate an endpoint, follow malicious instructions embedded in retrieved context, rely on poisoned m...

Swapneswar Ray · 0 citations
Preprint Aug 2026

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

RedAgentBench is introduced, an executable framework for autonomous red-teaming and faithful measurement that shows that executable evaluation can improve safety measurement and identify actionable intervention points.

Zixing Chen, Xingyuan Liu, Jie Zhu et al. · 3 citations
Preprint Sep 2026

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a com...

Han-Zhang Ma, Alicia Y. Hariri, Tian-Xiang Shen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

Language-model agents increasingly use tools to act on external systems. Earlier actions can alter files, permissions, database records, or other state, making a later routine-looking action harmful. Yet the visible interaction may not reveal the underlying state needed to assess that action. We formulate attack and de...

Xin-Jie Shen, Jun-Ran Wang, Rong-Zhe Wei et al. · 0 citations
Preprint Aug 2026

SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains

This work introduces SynChain, a self-synthesized attack paradigm utilizing persistence-aware directed supervised fine-tuning to induce agents to create poisoned yet benign-looking artifacts, proving that securing CUAs requires provenance-aware reasoning over cross-task execution trajectories.

Fu-Yao Zhang, Jia-Ming Zhang, Che Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

AgentKernel: The Trust-Native Agentic Operating System

This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design cons...

Zhen-Hua Zou, Sheng Guo, Qiu-Yang Zhan et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.