Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail t...
Jeremy Qin, David Schmotz, Derck W. E. Prinzhorn et al.· 0 citations
It is found that GPT-6 Astra's low evasion rate comes with overrefusal, as it frequently abandons otherwise solvable tasks under a denial-of-service prompt injection.
David Schmotz, Derck W. E. Prinzhorn, Luca Beurer-Kellner et al.· 1 citation
This work investigates an alternative standard designed to function despite ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for...
Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert· The Paris Journal on AI &...· 1 citation
Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, it is shown no model expresses a coherent policy across the three deployments, suggesting LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.
Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans· The Paris Journal on AI &...· 1 citation
ResearchArena is released as a modular framework for evaluating sabotage and control in automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization.
Lena Libon, Ben Rank, Jehyeok Yeon et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.