This paper analyzes how AI could persuade humans in key settings toward decisions that compromise the development, containment, oversight, and governance of AI itself, and provides a blueprint for assessing the associated risks.
Abstract
The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.
A risk-differentiated verification framework is developed that calibrates verification practice to the epistemic stakes of different AI uses, and argues that the property at stake is warrant: whether content has passed through a process that entitles a reader to rely on it.
A. Ünal· Ethics and Information Techn...· 0 citations
To continue exposing AI-driven military violence, OSI practices must reorient their methodologies and focus more systemically and systematically on the dispersed material infrastructures and political forces underpinning AI militarism, so investigators can better reveal concealed networks of accountability linking stat...
It is found that at least for the time being, explicit normative instructions are not fully able to realign AI advice with the normative convictions of the population, or the legislator deciding on its behalf.
Debates on catastrophic artificial intelligence (AI) risk often frame artificial superintelligence as a problem of value alignment: the central task is to ensure that advanced systems act in accordance with human intentions or moral principles. This article argues that such a framing is necessary but insufficient. The...
C. Burelli, Federico Formentini· European Journal of Politica...· 0 citations
This article advances a novel approach to the question of who should take responsibility when no one is to blame, and shows how it can be applied to the emerging problem of faultless harm caused by artificial intelligence (AI). As AI systems increase in complexity and scale, so too does the risk of harm that cannot b...
This paper uses disparate examples of AI resistance and theoretical lenses to render explicit an underlying logic which connects varied forms of resistance, grounded in the ethical value(s) that effort can manifest, and reframe one side of the debate between embracers and resisters.
R. Downes, Rebecca Mines· AI and Ethics· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.