PersonalBench is introduced, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics, which finds that personalization methods do produce author-differentiated output but this differentiation never crosses the human-LLM boundary.
Abstract
Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's writing. We introduce PersonalBench, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics. Across 50 authors, 1,000 generations, and two model families (Qwen 3, GLM-4), we find that personalization methods do produce author-differentiated output (LUAR discriminates target authors within generated text at AUC=0.918) but this differentiation never crosses the human-LLM boundary. All methods achieve LUAR similarity to real authors in the range 0.484-0.508, below the cross-author human floor of 0.626 (ceiling 0.756). The LLM's own authorship fingerprint dominates: generated text is more distant from any human author than random humans are from each other. Methods are statistically indistinguishable from each other on LUAR (spread 0.024) despite appearing differentiated on the LLM judge, a discrepancy we trace to circularity between trait extraction and profile extraction. We validate that LUAR reliably measures authorship in our corpus (AUC=0.76 single-post, 0.96 multi-post). We release PersonalBench as a calibrated measuring stick: inference-time personalization modulates the LLM's style but does not bridge the gap to human authorship.
AVShift is introduced, the first German benchmark for systematically evaluating AV under multiple distribution shifts, and feature analysis reveals stylistic features that remain stable across genres, while their relative importance varies depending on the specific genre transition.
Lotta Kiefer, Brisca Balthes, Christoph Leiter et al.· 0 citations
Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this...
Hao-Han Yuan, Si-Min Chen, Xi Niu et al.· 0 citations
AuthBench is introduced, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic, and position it not only as a new benchmark, but as a diagnostic resource for studying when and why authorship representations succeed or fail.
Mao-Xun Huang, Zhen-Xing Zhang, C. Cardie· 0 citations
Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identification. This paper invest...
A corpus of arXiv authors with solo papers from the Mathematics Genealogy Project graph is built, giving 5 total authors and ground-truth advisor-student pairings, where advisors sit closer in cosine distance to their students than a random same-field author does.
Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We...
Qian Ma, A. Squicciarini, Sarah Rajtmajer· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.