FakeMark is presented, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction, to motivate provenance-aware, multi-factor ownership protocols.
Abstract
Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark, a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction. Under a simplified linear decision-boundary model, targeted perturbations can acquire a nonzero component along the watermark-trigger direction; experiments on deep networks provide only conditional, setting-dependent support for this intuition. FakeMark caches selected convolutional and fully connected layer outputs from clean surrogate batches and injects them through stochastic multi-layer, channel-wise interpolation to improve transfer. Across 16 distinct architectures and an additional adversarially trained ResNet-50 checkpoint variant, over eight evaluated watermark variants, retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. ImageNet transfer varies substantially across checkpoint and surrogate settings, ranging from near zero to 0.99. Matched baselines and detector analyses motivate provenance-aware, multi-factor ownership protocols.
FeatMark is introduced, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro- features that remain natural to humans while providing a stronger, machine-verifiable provenance signal.
Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al.· 0 citations
Model watermarking is a commonly used ownership verification technique for protecting the copyright of deep learning models. However, in practically deployed black-box model service scenarios, existing methods typically rely on misclassification-based backdoor trigger mechanisms, or assume that the verifier can obtain...
Jing Xiao, Song Xiao, Chao Guo et al.· Journal of King Saud Univers...· 0 citations
A dual-path network is proposed to encode watermark information into both the generated image and the owner’s secret key, which achieves superior robustness against various adversarial attacks while maintaining high visual quality across diverse generative models.
Cong-Rong Li, Ling-Yun Yu, Pei-Qi Jiang et al.· Proceedings of the Thirty-Fi...· 0 citations
This work identifies two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction.
The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images while preserving visual plausibility, with concrete consequences for public discourse and vulnerable populations. We propose a robust semantic watermarking framework that...
With the increasing trend of open-sourcing deep neural network (DNN) models, protecting model ownership has become a critical challenge, particularly in high-stakes domains such as medical AI. Existing backdoor-based watermarking methods suffer from two key limitations: visible trigger patterns and the lack of a reliab...