Skip to content

I Am No One: Style-Aware Paraphrasing for Text Anonymization

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

This work proposes a style-aware, prompt-driven anonymization approach that uses pretrained large language models to construct compact stylistic profiles from minimal samples and rewrite text to suppress identifiable style markers while preserving meaning.

Abstract

Authorship attribution models can re-identify users from seemingly anonymized text by exploiting stable stylistic fingerprints, even after explicit identifiers are removed, posing a growing privacy risk for text publishing and analytics. This risk extends to speech-derived text such as ASR transcripts of meetings and call-center conversations, where stylometric leakage can persist even after acoustic anonymization. Differential privacy-based anonymization often severely degrades text quality and utility. We propose a style-aware, prompt-driven anonymization approach that uses pretrained large language models to construct compact stylistic profiles from minimal samples and rewrite text to suppress identifiable style markers while preserving meaning. Across blog and review datasets, our approach reduces authorship attribution F1 by 60-70% while maintaining content quality and readability, substantially outperforming DP-based and non-DP baselines.

View source

Similar papers

Preprint Aug 2026

Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We...

Qian Ma, A. Squicciarini, Sarah Rajtmajer · 0 citations
#natural language process... Preprint Aug 2026

Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation

Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identification. This paper invest...

Muhammed Nazmul Arefin, O. Hammad · 0 citations
#artificial intelligence Preprint Sep 2026

AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths

AuthBench is introduced, a large-scale multilingual benchmark for authorship representation that is designed to make this evaluation broad, standardized, and realistic, and position it not only as a new benchmark, but as a diagnostic resource for studying when and why authorship representations succeed or fail.

Mao-Xun Huang, Zhen-Xing Zhang, C. Cardie · 0 citations
Aug 2026

Emoji fingerprints: topic-independent authorship verification using 20 novel emoji-based stylometric features.

Authorship verification in social media text is a growing challenge in digital forensics, yet existing stylometric approaches rely on textual features that degrade substantially when reference and questioned texts originate from different communicative contexts. This study investigates whether emoji usage patterns-a pe...

Yasin Etli, Ahmet Özdil, M. Aşırdizer · 0 citations
#natural language process... Preprint Sep 2026

Forging LLM Authorship Fingerprints with Targeted Rewriting

Model-attribution classifiers can often identify which language model produced a text, making model-specific writing patterns a signal of provenance. Accurate attribution on unmodified text, however, does not show whether the prediction still identifies the original source after deliberate rewriting. We formulate this...

Hao-Han Yuan, Si-Min Chen, Xi Niu et al. · 0 citations
Review Open access Aug 2026

Recent Advances in Text Anonymization: A Systematic Review

This survey provides the first comprehensive and systematic review of text anonymization methods published between 2020 and 2025, covering 48 primary studies identified through a structured search and rigorous screening procedure and reveals a growing shift from identifier‐centric de‐identification toward context‐aware...

Marina Litvak, A. Jorge · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.