Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
These findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence, which positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.