Skip to content
Preprint

PersonalBench: Measuring the Authorship Gap in LLM Personalization

Aug 2026 · 1 citation · 20 references
Computer Science

TL;DR

PersonalBench is introduced, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics, which finds that personalization methods do produce author-differentiated output but this differentiation never crosses the human-LLM boundary.

Abstract

Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's writing. We introduce PersonalBench, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics. Across 50 authors, 1,000 generations, and two model families (Qwen 3, GLM-4), we find that personalization methods do produce author-differentiated output (LUAR discriminates target authors within generated text at AUC=0.918) but this differentiation never crosses the human-LLM boundary. All methods achieve LUAR similarity to real authors in the range 0.484-0.508, below the cross-author human floor of 0.626 (ceiling 0.756). The LLM's own authorship fingerprint dominates: generated text is more distant from any human author than random humans are from each other. Methods are statistically indistinguishable from each other on LUAR (spread 0.024) despite appearing differentiated on the LLM judge, a discrepancy we trace to circularity between trait extraction and profile extraction. We validate that LUAR reliably measures authorship in our corpus (AUC=0.76 single-post, 0.96 multi-post). We release PersonalBench as a calibrated measuring stick: inference-time personalization modulates the LLM's style but does not bridge the gap to human authorship.

View source

Similar papers

Preprint Aug 2026

When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era

AVShift is introduced, the first German benchmark for systematically evaluating AV under multiple distribution shifts, and feature analysis reveals stylistic features that remain stable across genres, while their relative importance varies depending on the specific genre transition.

Lotta Kiefer, Brisca Balthes, Christoph Leiter et al. · 0 citations
#natural language process... Preprint Sep 2026

Forging LLM Authorship Fingerprints with Targeted Rewriting

Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this...

Hao-Han Yuan, Si-Min Chen, Xi Niu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths

AuthBench is introduced, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic, and position it not only as a new benchmark, but as a diagnostic resource for studying when and why authorship representations succeed or fail.

Mao-Xun Huang, Zhen-Xing Zhang, C. Cardie · 0 citations
#natural language process... Preprint Aug 2026

Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation

Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identification. This paper invest...

Muhammed Nazmul Arefin, O. Hammad · 0 citations
#artificial intelligence Preprint Aug 2026

Writing Style Similarity Reflects Academic Genealogy

A corpus of arXiv authors with solo papers from the Mathematics Genealogy Project graph is built, giving 5 total authors and ground-truth advisor-student pairings, where advisors sit closer in cosine distance to their students than a random same-field author does.

Cameron Manzo · 0 citations
Preprint Aug 2026

Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We...

Qian Ma, A. Squicciarini, Sarah Rajtmajer · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.