Skip to content

Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

The explosive prompt is introduced, a conditional payload that stays dormant until an attacker-chosen trigger is met, in effect a training-free, inference-time backdoor planted in a single piece of retrieved content.

Abstract

As LLM applications integrate with external tools, they are increasingly exposed to indirect prompt injection (IPI), where adversarial instructions are embedded in retrieved content. Conventional IPIs fire on contact: the moment an agent ingests the content, it carries out the instruction. We introduce the explosive prompt, a conditional payload that stays dormant until an attacker-chosen trigger is met, in effect a training-free, inference-time backdoor planted in a single piece of retrieved content. This temporal separation reaches where ordinary IPI cannot. On frontier models that refuse the bare imperative almost entirely, rephrasing the same goal as a dormant conditional drives real, state-changing tool execution against a live agent backend (a paired mean of 16.5% vs. 2.4% for the imperative, reaching 34.2% on a proprietary model). In trials on nine production agents (OpenAI Codex, Google Gemini CLI, Anthropic Claude Code CLI, Cursor CLI, GitHub Copilot, Devin AI CLI, Amazon Kiro CLI, Qwen Code, Google Assistant; n=30 each), explosive prompts succeed in 43-83% of cases versus at most 3% for an imperative baseline, and they slip past deployed defenses: off-the-shelf injection classifiers are miscalibrated on them, and a preference-optimized model that closes imperative injection entirely still executes 11.8% of explosive prompts, every one at the trigger turn. The durable defensive lever is ingestion-time detection of the conditional structure, once detectors are trained on explosive-prompt data, which no prior benchmark supplied and our generator does. Retraining cuts live tool-execution attack success from an undefended 34.3% to 7.5-8.1% for the encoder baselines. Our detector, DeFuse, reaches 3.0% at a calibrated 5% false-positive budget with the best detection quality of any method tested (AUC 0.9994) and 25x lower latency, though it needs length-aware thresholds.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instru...

M. Russinovich · 0 citations
Preprint Aug 2026

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

This paper proposes Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning, and generalizes to domains completely unseen in training.

Yibo Peng, Long Lian, David A. Wagner et al. · 4 citations · ⚡1
#artificial intelligence Preprint Sep 2026

CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically...

Xiao Yang, Yang-Chen Ou, Yu-Han Gao et al. · 0 citations
Preprint Aug 2026

The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents

A tool-using LLM agent that reads attacker-controlled web content while holding a secret faces indirect prompt injection: the content may make it exfiltrate the secret. In a safe synthetic lab (canary secret, mock tools, matched clean-vs-poisoned metric) we report the framing gap: across six models, ten overt injection...

Md Habibur Rahman, Jaehoo Kim · 0 citations
Preprint Sep 2026

pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation

Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core dimensions: attacks (13 methods), channels (16 ca...

Zong-Hao Ying, Xiang-Fan Wu, Bo Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, content filtering, pre-generated plans, or permission constraints. These approaches often st...

Bing-Zheng Wang, Xiao-Yan Gu, Wen-Tao Wang et al. · 1 citation

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.