Skip to content
#small language model Conference Open access

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

Sep 2026 · Pacific-Asia Conference on Knowledge Discovery and Data Mining · pp. 240-255 · 1 citation · 32 references
Computer Science

TL;DR

CLIPGuard is proposed, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP.

Abstract

Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git

Read PDF

Similar papers

Preprint Sep 2026

Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers

TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples, is proposed.

A. Abdel-Naby, Mohamed Elmahallawy · 0 citations
Open access Sep 2026

PatchGuard-Freq: Zero-Overhead Adversarial Patch Defense via Frequency Detection and Data-Driven Robustness

Adversarial patch attacks pose a tangible physical-world threat to traffic sign recognition in autonomous driving systems. Current state-of-the-art defenses based on image reconstruction require dual-model deployment and add per-frame inference latency, making them impractical for resource-constrained embedded platform...

De-Jie Luan, Cheng-Hua Li, Chun-Jie Zhang et al. · 0 citations
Preprint Sep 2026

MROP: Mask-Region Optimized Purification Against Backdoor Attack in Deep JSCC

Deep joint source and channel coding (JSCC) transmits a source by mapping it directly to channel symbols through an end-to-end deep neural network (DNN) and reconstructing it at the receiver. Taking image transmission as an application, this DNN pipeline behaves as a black box: the receiver cannot readily detect securi...

Seongkyu Yang, Hyeonho Noh, Hyun Jong Yang et al. · 0 citations
Preprint Sep 2026

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

FeatMark is introduced, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro- features that remain natural to humans while providing a stronger, machine-verifiable provenance signal.

Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions and search for projection-matrix edits that suppress safety alignment while preserving general capabilities, within minutes on a single GPU and without gradient-based tra...

Tian Gao, Zhi-Hui Xie, Yu-Hao Wu et al. · 0 citations
Preprint Aug 2026

Adversarial Attacks on Deep OCR Systems

Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black...

Wenbo Sun, Hong-Zong Li, Yanyun Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.