Skip to content
Review Open access

Recent Advances in Text Anonymization: A Systematic Review

Aug 2026 · WIREs Data Mining and Knowledge Discovery · 0 citations · 57 references

TL;DR

This survey provides the first comprehensive and systematic review of text anonymization methods published between 2020 and 2025, covering 48 primary studies identified through a structured search and rigorous screening procedure and reveals a growing shift from identifier‐centric de‐identification toward context‐aware anonymization.

Abstract

Text anonymization has become a critical requirement across healthcare, legal, financial, and online communication domains, where large volumes of sensitive textual data are increasingly used for analytics, information retrieval, and model training. Despite decades of research, anonymization remains challenging due to the diversity of sensitive information types, domain‐specific annotation schemes, and the emergence of complex inference risks that extend beyond explicit identifiers. On the other hand, we have seen a large number of approaches using transformers in recent years. This survey provides the first comprehensive and systematic review of text anonymization methods published between 2020 and 2025, covering 48 primary studies identified through a structured search and rigorous screening procedure. We analyze anonymization approaches across rule‐based, statistical, neural, transformer‐based, and large language model (LLM) paradigms, and compare them across multiple domains, languages, and sensitive‐information taxonomies. Our review also synthesizes available datasets, software resources, and evaluation practices used to assess both privacy protection and text utility. The findings highlight significant fragmentation across domains, persistent limitations in handling quasi‐identifiers and semantic leakage, and substantial inconsistencies in evaluation protocols. We identify key methodological trends, gaps, and emerging challenges, including the integration of LLMs, multilingual settings, and adversarial evaluation. Our analysis reveals a growing shift from identifier‐centric de‐identification toward context‐aware anonymization, while evaluation methodologies have not yet evolved at the same pace. We outline open research directions and propose a roadmap toward more robust, context‐aware, and empirically grounded anonymization systems. This survey aims to establish a unified reference point for researchers and practitioners working on privacy‐preserving NLP and sensitive text processing.

Read PDF

Similar papers

Review Open access Aug 2026

Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions

Large Language Models (LLMs) have emerged as powerful tools that impact information integrity on social media platforms. This comprehensive review examines the dual role of LLMs in both facilitating and mitigating various information integrity challenges, including misinformation, disinformation, fake news, social bots, and privacy concerns. We conduct a comprehensive review of the literature from 2019 to 2024, screening 1048 studies and performing an in-depth analysis of 215 representative papers. This systematic approach allows us to identify key patterns in how LLMs influence the information security in social media ecosystems. Through a systematic analysis of papers from multiple databases, our findings reveal that while LLMs can enhance detection capabilities for malicious content and enable sophisticated defense mechanisms, they simultaneously pose risks by enabling the generation of highly convincing, deceptive content. We categorize and analyze the potential and challenges across different dimensions of information integrity, examining technical capabilities, ethical implications, and privacy concerns. The study demonstrates critical gaps in current approaches, particularly in cross-lingual detection, real-time monitoring, and privacy-preserving implementations. We conclude by proposing future research directions and recommendations for stakeholders to leverage LLMs while mitigating risks in social media information integrity.

Junjie Xiong, Zhengyuan Jiang, Xiaoran Xu et al. · 0 citations
Book Open access Aug 2026

Sensitive Data Detection in Documents with LLMs

Detecting and extracting sensitive information from documents is essential for privacy and regulatory compliance. Existing approaches either require training on large labeled datasets or rely on brittle, costly-to-curate pattern matching, while Large Language Models (LLMs) offer a promising alternative. We present a systematic evaluation of several proprietary and open-source LLMs for sensitive entity extraction from documents. Because large datasets of completed forms containing personal information are unavailable, we also introduce a form-filling pipeline that uses vision-capable LLMs to label form fields, generate realistic synthetic personas, and fill real blank forms, enabling reproducible evaluation across diverse layouts. Evaluating on these forms and the public RVL-CDIP dataset, we find performance is uneven across entity types—and that LLMs fall short of a simple pattern-based baseline on Social Security numbers.

Errita Xu, Stefan Larson, Kevin Leach · 0 citations
Jul 2026

Can We Explain What We Anonymize? On the Impact of Data Anonymization on Post-hoc Model Explanations

Privacy-preserving data publishing and explainable artificial intelligence (XAI) are both essential for trustworthy machine learning, yet their interaction remains largely underexplored. In practice, models are often trained on anonymized datasets, but little is known about how classical anonymization techniques affect post-hoc explanations. In this paper, we provide a systematic empirical study of how feature attribution rankings change under widely used anonymization models, including k-anonymity, $\ell$-diversity, t closeness, and $(\alpha, k)$-anonymity. Across multiple real-world datasets and classifiers, we compare explanations generated by SHAP and LIME and quantify their stability using rank correlation and hypothesis testing. Our findings reveal a fundamental trade-off: explainable privacy-preserving models are feasible under mild privacy constraints, but strict anonymization requirements often lead to unstable explanations and severe utility degradation.

Casper Lauge Nørup Koch, Mina Alishahi, Gaurav Choudhary · 1 citation

Future Generation Computer Systems

Jamila Alsayed Kassem, Tim Müller, Christopher A. Esterhuyse et al. · 0 citations