"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused
Results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents, and introduces CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks.