Skip to content
Preprint

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Aug 2026 · 0 citations
Computer Science

TL;DR

These findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence, which positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.

Abstract

Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.

View source

Similar papers

#natural language process... Preprint Aug 2026

Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

A systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view is conducted, highlighting the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse.

Lin Chen, Yi-Tong Chen, Yong Li · 0 citations
Preprint Apr 2026

When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation

The results suggest that directly asking for the target response may not always yield the most effective score for predicting it, and that comparing direct scores with indirect paths through related judgments may reveal a more effective predictive route.

Zonghuan Xu, Xiang Zheng, Yu-Tao Wu et al. · 0 citations
Open access Jul 2026

A Research Agenda for Studying LLM-Powered Persuasion Technology to Mitigate Misinformation

Emerging research on the use of large language models (LLMs) to counter false beliefs, including conspiracy theories, vaccine hesitancy, and climate denial, has relied predominantly on experimental methods rooted in social psychology, operationalizing persuasion as short-term attitudinal change measured through randomized controlled trials. We argue that this approach, while offering valuable causal inference, is insufficient for understanding how LLM-powered persuasion actually works. Drawing on communication theory and a qualitative analysis of publicly available transcripts from the “DebunkBot study” (Costello et al., 2024), we identify four constructs largely absent from existing research: users’ mental models of the LLM communicator, anthropomorphism (including negative forms such as attributions of naivety or gullibility that can drive resistance), folk theories about how AI systems operate, and epistemic literacy regarding machine-generated knowledge claims. We outline a research agenda organized around three priorities: methodological pluralism suited to the dynamic and personalized nature of LLM interactions, transparency and reproducibility standards for prompt design and transcript analysis, and ethical frameworks for assessing potential harms, including belief reinforcement among resistant participants and privacy risks in large-scale transcript data.

Claire Wardle, Pete Brown, David Scales · 0 citations
Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and steering of agent honesty.

Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan et al. · 0 citations
#small language model Preprint Aug 2026

Belief Cascades Drive Persuasion in LLM Agent Networks

This work introduces a controlled testbed for studying how goal-directed persuaders shift elicited stances in networks of LLM agents grounded in real-world ego-network topologies, and argues for evaluating multi-agent persuasion as a trajectory- and exposure-level process.

Haoyi Qiu, Genglin Liu, P. Venkit et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.