Skip to content

CRFormer: Multi-scale contrastive regularization transformer for visible watermark removal.

Aug 2026 · Neural Networks · Vol 205 Pt B, pp. 109504 · 0 citations · 40 references
Medicine

TL;DR

This paper proposes CRFormer, a single-stage Transformer network for blind visible watermark removal that replaces the CNN backbone with a full Transformer to model watermark regions of arbitrary spatial extent and introduces a deformable convolution feed-forward network that restores spatial perception and integrates watermark mask prediction directly into the backbone.

Abstract

Visible watermarks are widely used for image copyright protection, but their removal remains a challenging restoration task due to the diversity of watermarks in color, scale, transparency, and spatial distribution. Existing methods predominantly rely on CNN-based frameworks, where limited receptive fields constrain spatial modeling, and contrastive learning is applied at intermediate feature levels rather than directly supervising the final reconstructed output. In this paper, we propose CRFormer, a single-stage Transformer network for blind visible watermark removal. CRFormer replaces the CNN backbone with a full Transformer to model watermark regions of arbitrary spatial extent. To compensate for the reduced spatial sensitivity of channel-wise attention, we introduce a deformable convolution feed-forward network that restores spatial perception and integrates watermark mask prediction directly into the backbone. We further apply contrastive learning as an output-level regularization, where multi-scale VGG features of the reconstructed image are pulled toward the watermark-free ground truth and pushed away from the watermarked input, providing direct supervision over perceptual reconstruction quality that intermediate-feature decoupling cannot offer. Extensive experiments on multiple public benchmarks demonstrate that CRFormer consistently outperforms existing state-of-the-art methods by a significant margin.

View source

Similar papers

2026

Laplacian Pyramid Reweighting With Progressive Residual Learning for Image Forgery Localization

The increasing realism of image manipulations poses significant challenges for forgery localization. However, existing methods are hindered by the limited adaptability of constrained frequency filters and the dilution of subtle forensic cues in deep networks. To address these challenges, we propose the Laplacian pyramid reweighting with progressive Residual Learning framework (LapRL-Net). First, a Laplacian Residual Adaptive Reweighting (LRAR) module is introduced to adaptively modulate multi-scale frequency residuals, enabling flexible extraction of discriminative frequency artifacts. Second, to mitigate feature dilution, we design a Progressive global-local Residual Fusion Module (PRFM) with multi-level residual fusion, which progressively combines global contextual dependencies with local texture details to preserve critical forensic cues. Furthermore, an Edge-Guided Refinement Module (EGRM) is incorporated to enhance boundary accuracy by enforcing geometric consistency via edge supervision. Extensive experiments on multiple benchmarks demonstrate that the proposed method achieves competitive performance in complex forensic scenarios.

Zhuo-Fei Liu, Wen-Jie Li, Yang Yu · 0 citations
Jul 2026

SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack

Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capture the long-range dependencies of globally distributed watermark signals and resulting in an unfavorable trade-off between removal effectiveness and visual fidelity. In this paper, we propose SPFM-Net, a semantic-prior-guided and frequency-constrained Mamba framework for invisible watermark attack. SPFM-Net first employs high-ratio masking to disrupt the spatial coherence of invisible watermark signals, and then utilizes a partially fine-tuned pretrained Masked Autoencoder to reconstruct semantically consistent image from sparse observations while suppressing watermark-related information. A Multi-scale Residual Frequency Feature Interaction module subsequently aggregates watermark-related residual features across multiple receptive fields, while adaptively suppressing responses from watermark-irrelevant regions. To further capture the long-range dependencies of globally distributed watermark signals, a lightweight Mamba-based Global State-space Feature Modeling (GSFM) unit is introduced to separate watermark-related features from natural image content and suppress the remaining watermark traces. In addition, SPFM-Net is optimized using a multi-level objective that jointly imposes spatial-, frequency-, and edge-domain constraints, enabling effective watermark suppression while preserving perceptual quality. Extensive experiments on representative spatial-domain, transform-domain, orthogonal moment-based, and deep learning-based watermarking schemes demonstrate that SPFM-Net achieves a favorable trade-off between watermark attack effectiveness and perceptual fidelity.

Chun-peng Wang, Yan Shi, Zhi-qiu Xia et al. · 0 citations
Preprint Aug 2026

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

FIT employs a lightweight Degradation Encoder to predict a global degradation vector and a spatial degradation map from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation, and introduces a task-token dropout strategy that regularizes task conditioning during training.

Zihao He, Yunfeng Wu, Xinchao Wang et al. · 0 citations
Preprint Aug 2026

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore is presented, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining.

Lingchen Sun, Rongyuan Wu, Xiangtao Kong et al. · 0 citations
Jul 2026

FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack

Existing invisible watermark removal methods often struggle to accurately capture the watermark-bearing features, leading to an unfavorable trade-off between watermark suppression and perceptual fidelity. In this paper, we propose the Frequency-Decoupled Diffusion Watermark Attack Network (FDDWAN), a coarse-to-fine framework that performs watermark removal through wavelet-domain decomposition and residual diffusion refinement. In the initial stage, the Wavelet-based Frequency-domain Preliminary Attack Module (WFPAM) decomposes the watermarked image into low- and high-frequency subbands and applies frequency-specific attack strategies tailored to their respective contributions to watermark robustness and perceptual quality. In the next stage, the Frequency-domain Residual Diffusion Attack Module (FRDAM) separately models the residual distributions between the preliminarily attacked outputs and the corresponding watermark-free references during training. Rather than reconstructing the entire image, FRDAM selectively refines frequency-domain residuals, directing the diffusion process toward the remaining watermark related discrepancies while minimizing modifications to image content. Extensive experiments on CelebA and ImageNet across four representative watermarking schemes demonstrate that FDDWAN achieves a more favorable trade-off between watermark removal effectiveness and visual fidelity than conventional and learning-based attack methods.

Chun-peng Wang, Yuxin Li, Xiaoyu Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.