#artificial intelligence
May 2026
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
This work introduces AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents, and uncovers the Safe Source Paradox, a safety-utility trade-off for retrieval-enabled agents.
Aditya Nawal, Manit Baser, M. Gurusamy
· arXiv.org · 0 citations