Skip to content

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

Jul 2026 · arXiv.org · Vol abs/2607.13162 · 0 citations · 59 references
Computer Science

TL;DR

This work presents the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural, steerable latent but amplifiable, or intractable (resistant to standard extraction).

Abstract

What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like"evil,"a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.

View source

Similar papers

Jul 2026

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile by extracting activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale.

Hasibur Rahman, Smit Desai · 0 citations
Jul 2026

From Minds to Models: The Intersection of Psychology and LLM Behaviours

Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building...

Oliver A. Guidetti, Reza Ryan · 0 citations
Preprint Aug 2026

Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference

This work proposes PRISM (Persona Reasoning with Inverse SFL-based Modeling), a psycholinguistically grounded framework that reformulates persona fidelity evaluation as a structured inverse inference task, providing a more reliable framework for persona fidelity evaluation.

Meng-Fan Li, Ze-Sheng Wei, Xuan-Hua Shi et al. · 1 citation
Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses ac...

Katherine Van Koevering, Anjalie Field · 0 citations
#artificial intelligence Preprint Jul 2026

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even impr...

Sunny Rai, Jin-Yi Kuang, Reyhan Jamalova et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.