PsychJail is presented, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion and establishes psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.
Abstract
Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy. It factorizes each attacker action into a Change-of-Meaning analysis, tactic selection, and victim-visible message, operationalizing the Persuasion Knowledge Model (PKM). The policy is refined with trajectory-level reinforcement learning using a PKM-gated reward that credits early jailbreak success only when every turn contains a well-formed Change-of-Meaning analysis. Across four aligned victim models, PsychJail achieves the highest average attack success rate (87.3%) and outperforms strong single-turn and multi-turn baselines on every model. We also measure susceptibility at the action that breaks each victim, revealing four distinct model-level fingerprints that identify which persuasion levers affect each model and how broadly. These fingerprints help explain cross-model transfer asymmetry. We interpret them as four candidate psychological profiles-rationalist, credibility-driven, narrative-monoculture, and broadly persuadable-while treating this interpretation as a conjecture requiring future validation. Our findings establish psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.
BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.
Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al.· 0 citations
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified pro...
P. Fonseca, R. Rodríguez-Carvajal, Rafael A. Calvo· 0 citations
Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users'expressed views. In reality, sycophancy rarely happens in a single exchange; it m...
Sidharth Pulipaka, R. Binkytė, Ivaxi Sheth et al.· 0 citations
Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campa...
Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos et al.· 0 citations
Algorithmically curated social media feeds can provoke perceived anxiety in everyday use, yet little is known about how AI can provide effective support in these contexts. We present TriggerDetector, an LLM-powered system that infers plausible anxiety triggers from user-encountered posts and offers coping suggestions....
Yi-Wen Shang, Yang-Tao Ge, Lucy J. Robinson et al.· Proceedings of the 14th Nord...· 0 citations
Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind...
Siddhant Jain, Dimitra Tsovaltzi· Companion Publication of the...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.