TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning
LoRAs are widely studied for adapting base text-to-image diffusion models. However, a backdoored LoRA can hide a backdoor: it behaves normally in most cases, but produces attacker-specified content (the backdoor target) when a hidden backdoor trigger appears in the input prompt. Detecting such backdoors before using an...