Skip to content
Open access

FakeMark: gradient-guided false watermark claims via robust feature fusion

Sep 2026 · Cybersecurity · Vol 9 · 0 citations · 45 references

TL;DR

FakeMark is presented, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction, to motivate provenance-aware, multi-factor ownership protocols.

Abstract

Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction. Under a simplified linear decision-boundary model, targeted perturbations can acquire a nonzero component along the watermark-trigger direction; experiments on deep networks provide only conditional, setting-dependent support for this intuition. FakeMark caches selected convolutional and fully connected layer outputs from clean surrogate batches and injects them through stochastic multi-layer, channel-wise interpolation to improve transfer. Across 16 distinct architectures and an additional adversarially trained ResNet-50 checkpoint variant, over eight evaluated watermark variants, retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. ImageNet transfer varies substantially across checkpoint and surrogate settings, ranging from near zero to 0.99. Matched baselines and detector analyses motivate provenance-aware, multi-factor ownership protocols.

Read PDF

Similar papers

Preprint Sep 2026

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

FeatMark is introduced, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro- features that remain natural to humans while providing a stronger, machine-verifiable provenance signal.

Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al. · 0 citations
Open access Aug 2026

Backdoor-free model watermarking via differential verification of adversarial samples and multi-bit information extraction

Model watermarking is a commonly used ownership verification technique for protecting the copyright of deep learning models. However, in practically deployed black-box model service scenarios, existing methods typically rely on misclassification-based backdoor trigger mechanisms, or assume that the verifier can obtain...

Jing Xiao, Song Xiao, Chao Guo et al. · 0 citations
Conference Open access Sep 2026

Latents-Inv:Robust Semantic Watermark via Dual-Path Mutual Information Redundancy for Diffusion Models

A dual-path network is proposed to encode watermark information into both the generated image and the owner’s secret key, which achieves superior robustness against various adversarial attacks while maintaining high visual quality across diverse generative models.

Cong-Rong Li, Ling-Yun Yu, Pei-Qi Jiang et al. · 0 citations
Preprint Sep 2026

Semantic Watermarking for Malicious Image Manipulation Detection

The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images while preserving visual plausibility, with concrete consequences for public discourse and vulnerable populations. We propose a robust semantic watermarking framework that...

Yoonseo Kim, S. Baek, Jun-Yong Park · 0 citations
Open access 2026

Lattice-Quantization Identity Watermarking for Deep Neural Network Ownership Protection

With the increasing trend of open-sourcing deep neural network (DNN) models, protecting model ownership has become a critical challenge, particularly in high-stakes domains such as medical AI. Existing backdoor-based watermarking methods suffer from two key limitations: visible trigger patterns and the lack of a reliab...

Ying Xu, Zhi-Ying Li, Shanxiang Lyu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.