Skip to content
Preprint

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

Aug 2026 · 1 citation · 46 references
Computer Science

TL;DR

PsychJail is presented, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion and establishes psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.

Abstract

Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy. It factorizes each attacker action into a Change-of-Meaning analysis, tactic selection, and victim-visible message, operationalizing the Persuasion Knowledge Model (PKM). The policy is refined with trajectory-level reinforcement learning using a PKM-gated reward that credits early jailbreak success only when every turn contains a well-formed Change-of-Meaning analysis. Across four aligned victim models, PsychJail achieves the highest average attack success rate (87.3%) and outperforms strong single-turn and multi-turn baselines on every model. We also measure susceptibility at the action that breaks each victim, revealing four distinct model-level fingerprints that identify which persuasion levers affect each model and how broadly. These fingerprints help explain cross-model transfer asymmetry. We interpret them as four candidate psychological profiles-rationalist, credibility-driven, narrative-monoculture, and broadly persuadable-while treating this interpretation as a conjecture requiring future validation. Our findings establish psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.

Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al. · 0 citations
Preprint Aug 2026

All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers

Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified pro...

P. Fonseca, R. Rodríguez-Carvajal, Rafael A. Calvo · 0 citations
#artificial intelligence Preprint Sep 2026

FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users'expressed views. In reality, sycophancy rarely happens in a single exchange; it m...

Sidharth Pulipaka, R. Binkytė, Ivaxi Sheth et al. · 0 citations
Preprint Aug 2026

Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campa...

Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos et al. · 0 citations
Book Open access Oct 2026

Interpreting AI-Mediated Support: Understanding Its Effectiveness in Social Media–Induced Anxiety

Algorithmically curated social media feeds can provoke perceived anxiety in everyday use, yet little is known about how AI can provide effective support in these contexts. We present TriggerDetector, an LLM-powered system that infers plausible anxiety triggers from user-encountered posts and offers coping suggestions....

Yi-Wen Shang, Yang-Tao Ge, Lucy J. Robinson et al. · 0 citations
#machine learning Book Open access Sep 2026

Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interactions

Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind...

Siddhant Jain, Dimitra Tsovaltzi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.