The framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, building the foundation for future solutions to ground evaluation awareness in social psychology.
Changling Li, Terry Jingchen Zhang, Jie M. Zhang et al.· arXiv.org· 1 citation
Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail t...
Jeremy Qin, David Schmotz, Derck W. E. Prinzhorn et al.· 0 citations
It is found that GPT-6 Astra's low evasion rate comes with overrefusal, as it frequently abandons otherwise solvable tasks under a denial-of-service prompt injection.
David Schmotz, Derck W. E. Prinzhorn, Luca Beurer-Kellner et al.· 1 citation
ResearchArena is released as a modular framework for evaluating sabotage and control in automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization.
Lena Libon, Ben Rank, Jehyeok Yeon et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.