PersonalBench: Measuring the Authorship Gap in LLM Personalization
PersonalBench is introduced, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics, which finds that personalization methods do produce author-differentiated output but this differen...