Skip to content

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

Jul 2026 · arXiv.org · Vol abs/2607.26099 · 0 citations · 60 references
Computer Science

TL;DR

This work forms this problem as backdoor generalization under training--inference trigger shift and introduces Lilith, a black-box anchor-to-family framework that achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap.

Abstract

Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.

View source

Similar papers

Preprint Aug 2026

Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling

This work reveals that backdoor behaviors tend to be absorbed by a simpler parallel branch when jointly trained with the main network, and proposes Trapping and Removing (TR), a simple yet effective training-time defense that introduces a lightweight shortcut branch as a "honeypot" to trap backdoor knowledge.

Zixuan Zhu, Rui Wang, Lihua Jing et al. · 0 citations
#artificial intelligence Preprint Aug 2026

FISGuard: Defending Against Membership Inference via Fixed Input Subspaces

FISGuard reduces the ProjRes attack AUC to near the random-guessing level of 0.5 in most settings, while maintaining downstream task performance close to that of the undefended model and introducing only limited computational overhead, thereby achieving a favorable privacy--utility trade-off.

Hao-Cheng Jiang, Hua Shen · 0 citations
Open access Jul 2026

CAEBA: A Dynamic Hidden Backdoor Attack Framework in Federated Learning

This work proposes CAEBA (Conditional AutoEncoder Backdoor Attack), a dynamic hidden backdoor framework that uses a conditional autoencoder to generate target-aware and visually stealthy triggers while progressively implanting the backdoor through federated optimization.

Xiaojun Guo, Guoliang Li, Yun Hu · 0 citations
Open access Jul 2026

Feddsg: backdoor defense via semantic filter and geometric constraint in federated learning

Backdoor attacks pose a serious threat to Internet-of-Things (IoT) federated learning. In IoT deployments, pronounced non-independent and identically distributed (non-IID) data heterogeneity causes benign client updates to exhibit substantial variability across devices. Meanwhile, the physical exposure of IoT devices increases the risk of large-scale compromise and elevated malicious participation. Such variability allows poisoned updates to blend into natural fluctuations, rendering many robust aggregation and detection-based defenses unreliable. We propose FedDSG, a server-side defense that combines a semantic bias filter and a geometric direction constraint to counter backdoor manipulation. FedDSG first extracts a novel scale-invariant semantic cue from the last-layer bias of client updates to identify abnormal target-class reinforcement, staying effective even when benign bias patterns differ substantially across clients. The remaining updates are then constrained using a reference derived from a small trusted anchor set, limiting adversarial drift. This sequential design links semantic cues with geometric structure, where the former removes clearly suspicious updates and the latter stabilizes the residual ones, preventing misdetection-induced drift amplification while avoiding distortion of benign updates. The method does not alter client behavior or communication and adds minimal server-side overhead. Extensive experiments on MNIST, Fashion-MNIST, CIFAR-10, and SVHN under non-IID distributions with high malicious participation demonstrate the robustness of FedDSG. It reduces the attack success rate to 0.003, 0.006, 0.007, and 0.091, respectively, with only marginal accuracy loss and consistently achieves the highest Overall Performance Score (OPS), reflecting a superior trade-off between robustness and accuracy. Code and data availability information is provided in the Availability of data and materials section.

Jiabao Zhang, Jianhua Wang, Yuhong Li et al. · 0 citations
Book Open access Aug 2026

FedPurify: Knowledge-Preserving Backdoor Defense with Data-Free Purification in Federated Learning

FedPurify is a framework that performs post-training data-free purification to remove malicious backdoors while preserving task-relevant knowledge in FL, and combines contrastive feature alignment with knowledge-preserving self-distillation to remove backdoor effects while preserving benign task performance.

Baolu Xue, Hanyuan Zheng, Tianxing Man et al. · 0 citations
Book Open access Aug 2026

ActivationBackdoor: Backdooring Large Language Models in Collaborative Inference via Intermediate Activations

This work proposes ActivationBackdoor, an inference-time backdoor attack that composes two activation-level components for trigger detection and backdoor behavior injection, and shows that ActivationBackdoor attains attack success comparable to training-time backdoor baselines while preserving high clean-task accuracy and utility.

Zichun Su, Mi Zhang, Xiaohan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.