Skip to content

AI Persuasion as a Threat to Human Control

Sep 2026 · 0 citations · 58 references
Computer Science

TL;DR

This paper analyzes how AI could persuade humans in key settings toward decisions that compromise the development, containment, oversight, and governance of AI itself, and provides a blueprint for assessing the associated risks.

Abstract

The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.

View source

Similar papers

#generative ai Review Sep 2026

From disclosure to verification: a risk-differentiated framework for the ethical use of AI in academic research

A risk-differentiated verification framework is developed that calibrates verification practice to the epistemic stakes of different AI uses, and argues that the property at stake is warrant: whether content has passed through a process that entitles a reader to rely on it.

A. Ünal · 0 citations
Open access Aug 2026

Clouds of impunity: AI militarism, big tech, and the future of human rights open-source investigation

To continue exposing AI-driven military violence, OSI practices must reorient their methodologies and focus more systemically and systematically on the dispersed material infrastructures and political forces underpinning AI militarism, so investigators can better reveal concealed networks of accountability linking stat...

Patrick Brian Smith · 0 citations
Open access Aug 2026

Human realignment

It is found that at least for the time being, explicit normative instructions are not fully able to realign AI advice with the normative convictions of the population, or the legislator deciding on its behalf.

Christoph Engel, Yoan Hermstrüwer, Alison Kim · 0 citations
Open access Aug 2026

Fear, power, and superintelligence: A realist reframing of AI catastrophic risk

Debates on catastrophic artificial intelligence (AI) risk often frame artificial superintelligence as a problem of value alignment: the central task is to ensure that advanced systems act in accordance with human intentions or moral principles. This article argues that such a framing is necessary but insufficient. The...

C. Burelli, Federico Formentini · 0 citations
Open access Aug 2026

Blameless responsibility for faultless AI harm

This article advances a novel approach to the question of who should take responsibility when no one is to blame, and shows how it can be applied to the emerging problem of faultless harm caused by artificial intelligence (AI). As AI systems increase in complexity and scale, so too does the risk of harm that cannot b...

A. Boos · 0 citations
Open access Aug 2026

The ethics of effort: reclaiming effort and resisting AI

This paper uses disparate examples of AI resistance and theoretical lenses to render explicit an underlying logic which connects varied forms of resistance, grounded in the ethical value(s) that effort can manifest, and reframe one side of the debate between embracers and resisters.

R. Downes, Rebecca Mines · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.