Skip to content
Preprint

Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples, is proposed.

Abstract

Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted misclassification while preserving high clean accuracy. Existing defenses often rely on model internals, training data, or clean validation samples, making them difficult to deploy when only black-box access to a trained model is available. We propose TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples. The key insight behind TRIM is to identify image regions that are responsible for anomalous model behavior and purify only those regions while preserving benign content. TRIM innovates via three key components: (i) region-based segmentation with deep feature representations, (ii) adaptive trigger discovery through inpainting and diffusion-based reconstruction to isolate regions responsible for misclassification---without assumptions about trigger type, shape, or location, and (iii) selective region purification that cleans poisoned regions while retaining benign content. To support practical deployment, TRIM further caches feature embeddings of previously identified triggers, enabling efficient recognition and avoiding redundant detection and purification. Extensive experiments across diverse datasets and backdoor types, including blended, sparse, varying-size, and multiple triggers, show that TRIM consistently outperforms existing black-box defenses, reducing attack success rates (ASR) to as low as 1.16% while preserving clean accuracy of up to 87.87%. These results demonstrate that effective backdoor mitigation is possible at inference time even when the defender has no access to any auxiliary data.

View source

Similar papers

Preprint Sep 2026

ODPure: Backdoor Purification for Object Detection via Ensemble Corruption Consensus

ODPure is proposed, a novel input-stage black-box defense for object detection, which is based on input purification that ensures stable perception flows and provides robust defense against diverse backdoor attacks and trigger types while preserving baseline accuracy.

Li Zeng, Ming-Cheng Duan, Long-Fei Fan et al. · 0 citations
Conference Open access Sep 2026

Mask-Guided Hybrid Triggers for Robust Clean-Label Backdoor Attacks

Clean-label backdoor attacks pose significant security threats to deep neural networks by injecting triggers without altering ground-truth labels. However, existing methods face a fundamental dilemma: sample-agnostic triggers are robust but easily detectable, while sample-specific triggers offer superior stealthiness b...

Shengye Pang, Xiang-Yu Ji, Jungang Yang et al. · 0 citations
#small language model Conference Open access Sep 2026

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

CLIPGuard is proposed, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP.

A. Abdel-Naby, Mohamed Elmahallawy · 1 citation
#artificial intelligence Preprint Aug 2026

Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift

DistScan is presented, a backdoor detection framework based on a simple but previously unexploited observation: backdoor injection systematically shifts a model's pre-NMS prediction class distribution away from its training class frequencies, even on clean inputs without any trigger present.

Longtian Wang, Zheng-Yu Zhao, Chen-Hao Lin et al. · 0 citations
Sep 2026

Seeing Through Threats: Adversarial Detection Through Explainability (ADEx)

Deep Neural Networks (DNNs) remain vulnerable to adversarial perturbations, raising significant concerns in image processing applications, particularly in high-stakes domains such as medical imaging and security-critical systems. Most existing defense strategies are limited by domain specificity, architectural dependen...

Syamantak Sarkar, Nirmal Joseph, Sudhish N. George et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.