#artificial intelligence
Oct 2025
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
The key findings show that while some detectors can identify attacks that rely on explicit textual instructions or visible image perturbations with moderate to high accuracy, they largely fail against attacks that omit explicit instructions or employ imperceptible perturbations.
Yinuo Liu, Ruohan Xu, Xilong Wang et al.
· arXiv.org · 22 citations
· ⚡1