Skip to content
Preprint

Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

DFCS is proposed, a training-free, trigger-agnostic method that clusters fixed pretrained features into one region per poisoning slot and selects the centroid-nearest sample from each region and supports distributional feature coverage as an effective selection principle for low-budget dirty-label backdoor attacks.

Abstract

Backdoor attacks compromise training data so that a model retains clean accuracy but predicts an attacker-chosen target on triggered inputs. At very low poisoning rates, only a few samples convey the trigger--target association, making poison-sample selection critical. Existing methods typically rank candidates using per-sample scores, which can select redundant samples from similar semantic regions, and many require task-specific surrogate training. We propose Distributional Feature Coverage Sample Selection (DFCS), a training-free, trigger-agnostic method that clusters fixed pretrained features into one region per poisoning slot and selects the centroid-nearest sample from each region. A local first-order analysis relates this allocation to feature-coverage and representative-mass terms. Across BadNets and Blended attacks on CIFAR-10, Tiny-ImageNet, and Imagenette, DFCS achieves the highest mean attack success rate among seven selectors in all six dataset--attack settings, averaging $96.30\%$ and exceeding the strongest comparator in each setting by 4.60 percentage points on average while preserving clean accuracy. These results support distributional feature coverage as an effective selection principle for low-budget dirty-label backdoor attacks.

View source

Similar papers

Jul 2026

Anti-Backdoor Coreset Selection via Cumulative Entropy

This paper formulate this defense strategy as a coreset selection problem, giving rise to so-called anti-Backdoor Coreset Selection, and uses the Cumulative Entropy as selection criterion to further facilitate this effect.

Qi Zhao, Christian Wressnegger · 0 citations
#machine learning Preprint Sep 2026

Empirical Evaluation of Data Poisoning Attacks in Supervised Learning

Data poisoning corrupts training data to degrade a model or to plant attacker-controlled behavior. This study evaluates two representative training-time attacks, label flipping and backdoor poisoning, on MNIST and Fashion-MNIST with three baseline classifiers: Logistic Regression, Linear SVM, and Random Forest. Clean training is compared with poisoning rates of 5%, 10%, and 20% using clean-test accuracy, macro-precision, macro-recall, macro-F1, and, for backdoors, attack success rate. Label flipping caused clear degradation, largest for Logistic Regression and Linear SVM, while Random Forest stayed comparatively stable. Backdoor poisoning reached attack success rates from 0.9667 to 1.0000 on both datasets and all three models while often keeping clean-test performance near baseline. The results separate indiscriminate poisoning, which shows up in standard metrics, from targeted backdoor poisoning, which stays comparatively stealthy while embedding highly effective malicious behavior, and they support security-oriented evaluation beyond conventional clean-test metrics.

Toshif Khan (Minot State University), Muhammad Abusaqer (Minot State University) · 0 citations
Jul 2026

Boundary Sampling for Efficient Model Extraction

This work proposes a novel data-free model extraction attack that substantially outperforms current methods in efficiency, accuracy, and overall effectiveness. Conventional black-box attacks depend heavily on treating the victim model as an oracle to label a large number of samples, primarily within high-confidence regions. This strategy not only demands an excessive number of queries but also often leads to the extraction of models with lower accuracy and limited transferability. In contrast, our method shifts focus to sampling low-confidence regions (along the decision boundaries) and leverages an evolutionary algorithm to enhance the sampling process. This approach dramatically reduces the query requirement by a factor of 10x to 600x, while also increasing the accuracy of the extracted model. Furthermore, our method achieves improved boundary alignment, significantly enhancing the transferability of adversarial examples from the extracted model to the victim, increasing the attack success rate from an average of 60% to 82%. Remarkably, these improvements are accomplished under a strict black-box scenario with soft-label (class-probability) query access, and no prior knowledge of the target model’s architecture or data distribution. Finally, we offer extensions to the algorithm to enable it to work on complex models: with high resolution, many classes, and even models with class imbalance such as anomaly detectors. Our attack is thoroughly evaluated on multiple image datasets with varying resolutions and is benchmarked against many state-of-the-art model extraction techniques. Additionally, to illustrate the versatility and robustness of our method, we conduct extensive experiments on four tabular datasets that vary in class numbers and sizes.

Doron Ben Chayim, Maor Biton Dor, Eyal Lenga et al. · 0 citations
#artificial intelligence Preprint Aug 2026

FISGuard: Defending Against Membership Inference via Fixed Input Subspaces

FISGuard reduces the ProjRes attack AUC to near the random-guessing level of 0.5 in most settings, while maintaining downstream task performance close to that of the undefended model and introducing only limited computational overhead, thereby achieving a favorable privacy--utility trade-off.

Hao-Cheng Jiang, Hua Shen · 0 citations

: Making

Politecnico Yuqi Zhao, I. di Torino, Politecnico Di et al. · 0 citations
Jul 2026

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

This work forms this problem as backdoor generalization under training--inference trigger shift and introduces Lilith, a black-box anchor-to-family framework that achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap.

Zhou Feng, Jia-Hao Chen, Chun-Yi Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.