Preprint
Aug 2026
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness, and ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.
Yutao Mou, Pengfei Yang, Zhenfei Yin et al.
· 0 citations