Deep neural networks (DNNs) are vulnerable to backdoor attacks, where the backdoored models behave normally on benign samples but misclassify trigger-carrying samples. However, when triggers are introduced as external cues inconsistent with original images, the resulting distribution shift makes existing backdoor attacks vulnerable to defenses based on abnormal latent representation detection. We propose Chameleon, a backdoor attack framework that reformulates sample-specific trigger generation as the restoration of a salient masked region. Chameleon combines diffusion-based image restoration guided by surrounding context with saliency-based mask positioning to generate semantically consistent triggers that are less separable from benign samples in latent space. We compare Chameleon with five baseline attacks on three datasets under six state-of-the-art backdoor defense methods, that is, STRIP, SentiNet, RNP, CCA-UD, SCAn, and Beatrix. Experimental results demonstrate that Chameleon achieves a 28.42% higher effective attack success rate (E-ASR) (i.e., successful attacks that evade defenses) compared to the best baseline when averaged across all defenses.
Boyang Zhou, Yixin He, Xiaofu Chen et al.· IEEE Transactions on Neural...· 0 citations
Text-to-image (T2I) models can be exploited to produce unsafe images. Existing safety measures, e.g., content moderation or model alignment, can be weakened by adversaries who attempt to restore unsafe generation through model fine-tuning. This paper presents Patronus, a defensive framework that improves T2I models’ resistance to the gradient-based adversarial fine-tuning attacks evaluated in this work. Specifically, we design a co-trained safety decoder that produces a deliberately corrupted output for a latent representation associated with unsafe content while preserving normal decoding for benign content. We also strengthen the decoder and U-Net with a non-fine-tunable learning mechanism. Across I2P, SneakyPrompt, and MMA-Diffusion, Patronus obtains attack success rates of 0.01–0.03 and true positive rates of 0.98–0.99. On benign prompts, it obtains FID 23.6, LPIPS 0.78, and a false positive rate of 0.01. The fine-tuning stress tests separately report the optimization losses of the defended decoder and U-Net under the evaluated attack settings.
Xinfeng Li, Sheng-Yuan Pang, Jialin Wu et al.· IEEE Transactions on Informa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.