Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents
The results show that the dominant source of failure can shift across model generations, motivating evaluation that diagnoses where and why long-horizon security agents fail rather than relying only on aggregate task success.