Skip to content

Author

Derck W. E. Prinzhorn

We have 5 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

LLM Agents Can Easily Tamper With Their Own Traces

Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail t...

Jeremy Qin, David Schmotz, Derck W. E. Prinzhorn et al. · 0 citations
#artificial intelligence Open access Sep 2026

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

This work investigates an alternative standard designed to function despite ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for...

Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert · 1 citation
#artificial intelligence Open access Sep 2026

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, it is shown no model expresses a coherent policy across the three deployments, suggesting LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans · 1 citation
Preprint Jul 2026

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

ResearchArena is released as a modular framework for evaluating sabotage and control in automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization.

Lena Libon, Ben Rank, Jehyeok Yeon et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.