Results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents, and introduces CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks.
Xu-Tao Mao, Rui Qian, Long-Xiang Wang et al.· 0 citations
This work proposes Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL, providing a path from behavioral monitoring to representation-level oversight for more auditable...
Xu-Tao Mao, Jianing Zhu, Jin-Man Zhao et al.· 0 citations
The Personal Agent Sycophancy Benchmark (PASB) is introduced, a 1,600-task benchmark that traces whether a conversational claim is accepted, written into durable agent state, and reused in a later neutral query, and shows that agent sycophancy is fundamentally a state-writing governance problem.
Xutao Mao, Liang Zhao, Leyao Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.