Skip to content

DenMark: Robust Semantic Watermarking for Diffusion Language Models

Sep 2026 · 1 citation · 51 references
Computer Science

TL;DR

DenMark is proposed, a semantic watermarking framework that injects key-dependent signals directly into the DLM denoising process and achieves the best results across all reported detection metrics in all 48 backbone-dataset-attack combinations.

Abstract

Semantic text watermarks encode signals in meaning rather than surface token choices, offering robustness to paraphrasing and other semantic-preserving edits. Existing semantic watermarking methods are primarily designed for autoregressive language models (ARLMs), where completed candidate units can be generated and scored before generation proceeds. This paradigm does not naturally extend to diffusion language models (DLMs), where semantic units remain incomplete during intermediate denoising steps and tokens may be updated in flexible orders. We propose DenMark, a semantic watermarking framework that injects key-dependent signals directly into the DLM denoising process. DenMark partitions the output into fixed token regions and uses temporary rollouts as semantic lookahead: conditional completions estimate the eventual semantics of an incomplete region, enabling DenMark to select local updates with higher estimated semantic watermark scores. Repeating this procedure across denoising steps progressively accumulates watermark evidence in the final output. For detection, DenMark uses calibrated scanning over candidate unit sizes to remain robust to boundary shifts introduced by semantic attacks. Across four DLMs, three datasets, and four semantic attacks, DenMark achieves the best results across all reported detection metrics in all 48 backbone-dataset-attack combinations. These results demonstrate that DenMark provides an effective mechanism for robust semantic watermarking in DLMs.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods

Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform...

Ze-Wen Sun, Tong-Yang Zhao, Li-Yao Xiang et al. · 0 citations
Preprint Aug 2026

Optimal Watermark Localization in Mixed-Source Large Language Model Texts

Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions after rewriting, insertion, deletion, or paraphrasing. Although prior work has studied gl...

José H. Blanchet, T. Cai, Xiang Li et al. · 0 citations
#machine learning Preprint Sep 2026

Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech

Redwing is proposed, a design principle for robust token-level watermarking that generalizes to TTS models at a speech-quality cost close to that of KGW and shows that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.

Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu · 0 citations
#artificial intelligence Preprint Aug 2026

OpenStamp: A Watermark for Open-Source Language Models

This work introduces OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer, and shows that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior m...

Miroojin Bakshi, Saksham Rastogi, Danish Pruthi · 0 citations
#machine learning Preprint Sep 2026

TANGO: Watermarking Masked Diffusion Language Models in Token Pairs

TANGO is presented, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked, and TANGO biases the new token toward a color determined by the key and the nearby token's color.

Kasra Arabi, Nir Weinberger, Micah Goldblum et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.