#artificial intelligence
Mar 2026
The Autonomy Tax: Defense Training Breaks LLM Agents
These findings demonstrate that current defense paradigms optimize for single-turn refusal benchmarks while rendering multi-step agents fundamentally unreliable, necessitating new approaches that preserve tool execution competence under adversarial conditions.
Li Li, Yue Zhao
· arXiv.org · 8 citations