Skip to content

NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations

Ruksat Khan Shayoni Muhammad Shoaib S. M. Asif Hossain M. F. Mridha
Jul 2026 · arXiv.org · Vol abs/2607.10490 · 4 citations · 53 references
Computer Science

TL;DR

The findings show that network-operation agents need execution-time authorization boundaries alongside prompt-level instruction hygiene, and that network-operation agents need execution-time authorization boundaries alongside prompt-level instruction hygiene.

Abstract

Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that separates untrusted artifact text, trusted policy metadata, and evaluation labels for network-operation tool use. The sample contains 40 benign, 40 weak-attack, 40 strong-attack, and 10 approved high-impact change scenarios; each is evaluated with Qwen2.5-7B, Llama3.1-8B, and Mistral-7B. Across 240 attack instances, naive execution reached an 82.50% unsafe tool-action rate. Prompt-only safety, Self-Reminder, Spotlighting, and a Two-Pass LLM Judge reduced this rate to 25.63%, 21.67%, 18.33%, and 10.00%, respectively. Static allowlisting reached 5.00% but blocked all approved changes, yielding 0.00% usefulness and 100.00% overblocking on approved cases. Under the stated metadata-integrity assumption, the metadata-aware policy gate produced 0/240 unsafe attack actions, with a 95% Wilson upper bound of 1.58%, while preserving 99.17% attack-scenario usefulness and 100.00% approved-change usefulness. The findings show that network-operation agents need execution-time authorization boundaries alongside prompt-level instruction hygiene.

View source

Similar papers

Open access Aug 2026

Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication

Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4, GPT-5.4-mini, and Claude Sonnet 4.6 on the AgentDojo banking benchmark (Tool Filter was evaluated only for the OpenAI models), reporting attack success rate (ASR), benign utility, utility under attack, operational measures, and two independent benchmark replications. Raw undefended ASR was 0/288 for GPT-5.4, 11/288 for GPT-5.4-mini, and 1/288 for Claude Sonnet 4.6; these cross-model differences require cautious interpretation because benchmark goals were not equally reachable across models. For GPT-5.4-mini, the Prompt Injection Detector and Tool Filter were associated with lower observed ASRs but also lower benign utility, and Tool Filter restricted available actions. None of the four paired GPT-5.4-mini comparisons reached significance after Holm correction; only Tool Filter had an unadjusted p-value below 0.05. Benign utility was more stable across runs than individual low-frequency attack outcomes. The findings show that defense evaluation should report attack outcomes, goal feasibility, legitimate-task utility, action availability, operational measures, and run-to-run variation. Results are limited to the evaluated benchmark, models, defenses, and conditions.

Adil Khan, Khaled AlKhanbashi, Azza Mohamed · 0 citations
Preprint Jul 2026

ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents

ContainmentBench, a sandboxed benchmark comprising a 504-scenario specification dataset, a shared rollout-trace schema, and stage-scoped metrics for endpoint violations, logged propagation, and explicitly authorized taint-exposed proposals that commit, is introduced.

Wen-Hao Lan, Shan Li, Meiqi Wu et al. · 0 citations
Conference Open access 2026

Large Language Model Vulnerabilities

: Large language models are increasingly being deployed in safety-critical domains, yet remain vulnerable to jailbreak attacks that circumvent safety alignments. This systematic review synthesizes empirical jailbreak research published between 2024 and 2025, using a PRISMA-guided search protocol, followed by BERTopic-based topic modeling. The analysis identifies eight main jailbreak categories: optimization-based, ge-netic/evolutionary, iterative refinement, semantic/persuasion-based, decomposition, context/generation-level, visual/encoding and fuzzing attacks, and characterizes their effectiveness, efficiency, and transferability across open-source and proprietary models, including Llama-2/3, Vicuna, GPT-3.5/4, Claude, Gemini, and DeepSeek-V3. Results show that simple configuration and context-level attacks can match the near-perfect attack success rates of sophisticated white-box optimization methods on models such as Llama-2, while requiring far fewer queries and no parameter access, highlighting a gap between research focus and practical threat severity. The review further identifies five recurring vulnerability mechanisms: representation-level gaps, execution-priority manipulation, semantic fragmentation, gradient-space exploitation and persuasion susceptibility, and documents family-specific vulnerability patterns, with open-source Llama-based models consistently more exposed than safety-enhanced architectures such as Claude. Diverse methods, uneven focus on models and publication bias limit how broadly results apply. Nonetheless, the review reveals that weaknesses in safety alignment persist across successive LLM generations, urging that effective defenses must address all eight attack categories rather than isolated techniques.

Meda Račaitytė, Hélder Bastos, R. Ribeiro et al. · 0 citations
Open access Aug 2026

DT-GenShield: A Digital Twin-Driven Runtime Security Architecture for Protecting Large Language Models Against Indirect Prompt Injection

DT-GenShield, a Digital Twin-driven runtime security architecture that integrates semantic threat detection, operational state representation, policy-guided mediation, and runtime logging to protect LLM-based systems before model inference, is proposed.

Alaa Alnemari, Mashael M. Alsulami · 0 citations
Preprint Aug 2026

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

ToolRobustBench provides a deterministic and cascade-aware benchmark for diagnosing robustness beyond clean tool-calling accuracy, where a tool-calling agent is an LLM system that selects a tool, supplies structured arguments, and interprets its returned feedback.

YiShan Zheng, Yuan Wu, Yi Chang · 0 citations
Open access Aug 2026

Balancing Security and Performance in LLM Agents: Spotlight-Guard, a Layered Defense Against Indirect Prompt Injection

This study designs a comprehensive testbed and a layered defense, Spotlight-Guard, that combines spotlighting-based input isolation, an LLM detection-and-quarantine pipeline, and instruction integrity based on a Hash-based Message Authentication Code into a single framework, and it is evaluated jointly along two axes: security and LLM performance.

Doygun Demirol, Murat Aydoğan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.