Jul 2026· International Conference on Smart Communications and Networking· pp. 1-6· 0 citations· 16 references
Abstract
Prompt injection poses a significant security risk to Retrieval-Augmented Generation (RAG) systems, enabling adversaries to embed malicious instructions within retrieved documents and hijack model behavior to exfiltrate sensitive information or execute unauthorized actions. This work presents a modular dynamic evaluation environment for systematically testing and comparing defense mechanisms against prompt injection attacks in RAG architectures. The framework simulates diverse injection scenarios targeting the retrieval pipeline, integrates optional mitigation strategies such as input filtering, prompt rewriting, and retrieval-aware defenses, and automatically logs model behavior to assess attack success. By varying retrieval parameters and quantifying defense robustness across diverse attack verticals and RAG configurations, the system enables reproducible and scalable evaluation of prompt injection resilience. The results highlight strengths and weaknesses of existing defenses in RAGspecific threat models and establish a foundation for standardized benchmarking of defense mechanisms in knowledge-grounded generative AI systems.
This work proposes Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware that matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines.
Zi Liang, XiaoYu Xu, Yanyun Wang et al.· 0 citations
A systematic review and structured descriptive synthesis of research on defenses against prompt-based attacks in language model and agent systems reveals trade-offs between security effectiveness, performance, and system complexity as well as major gaps in benchmarks, indirect attack coverage, and multi-agent evaluation.
Sana Mourad, E. E. Abdallah, Mohammad Ababneh· Electronics· 0 citations
This paper proposes SecureMCP, a policy-enforced framework that integrates Role-Based Access Control with an MCP server to establish multi-layer defense for LLM-generated SQL execution, and evaluates filter performance—false positive rate (FPR) and false negative rate (FNR))—separately from LLM generation quality.
Wonbae Kim, Hee-Kyong Yoo, Nammee Moon· Applied Sciences· 0 citations
The results establish that LLM-based log analysis creates an inherent confused deputy vulnerability where untrusted data and trusted instructions compete indistinguishably for model attention, requiring defense in-depth architectures and continued human oversight for security-critical decisions.
Rabimba Karanjai, Yang Lu, H. Madhavarao et al.· 2 citations
Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness, and ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.
Yutao Mou, Pengfei Yang, Zhenfei Yin et al.· 0 citations
Compositional Attack Path Scoring (CAPS), a framework engineered to quantify end-to-end multi-hop risks in LLM architectures, establishes a rigorous benchmark for quantitative vulnerability management in complex, agentic LLM environments.