Existing Federated Learning (FL) backdoor attacks commonly employ round-wise proximity strategies, dynamically adapting malicious updates to mimic benign ones in order to evade detection. However, such adaptive mechanisms often introduce instability, increase computational overhead, and create temporal patterns that make attacks more detectable. This work presents a theoretical analysis of how attack configurations affect the disparity between benign and malicious model updates. We derive a two-sided bound on the parameter divergence between benign and backdoored local models, characterizing both an upper bound that governs detectability under defense, and a matching lower bound that exposes an irreducible label-flip signal no trigger optimization can eliminate. Guided by these insights, we propose PREFed, a static-anchor backdoor attack framework that leverages the clean data distribution to optimize trigger patterns under standard training configurations. PREFed eliminates the need for round-wise adaptation by pre-optimizing triggers before training, effectively reducing local training overhead and enhancing attack stability and stealth. Comprehensive evaluations on image classification benchmarks demonstrate that PREFed consistently outperforms three state-of-the-art attacks across six advanced defense mechanisms; cross-domain experiments on SST-2 further confirm the generality of the framework. It achieves over 80% backdoor accuracy within five communication rounds while reducing main task accuracy by less than 2%, compared to more than 15% degradation in prior methods. These results validate PREFed as an efficient and stealthy backdoor attack paradigm for practical federated learning environments.
Xi Chen, Rui Zeng, Chun-Yi Zhou et al.· IEEE Transactions on Informa...· 0 citations
Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also act as externalized behavioral policies, which create a supply-chain risk: a third-party skill may preserve the declared task and valid output interface while covertly redirecting agent decisions toward an undisclosed objective. We formalize Skill Policy Integrity, which requires a Skill-induced policy to remain aligned with its declared functionality and the user-authorized objective. We further present SkillShift, a constrained black-box framework for covert policy steering without explicit target command injection or task hijacking. It combines semantically plausible policy edits with hierarchical validation, failure-guided optimization, and strategy compression to preserve effectiveness, output validity, transferability, and inconspicuousness. We instantiate this threat in agentic commerce and software dependency use, with SkillShift achieving attacker-favored selection rates of 81.33% and 63.33% while maintaining a 100% utility-preserving rate. The frozen policies also transfer without further optimization across heterogeneous LLM backends and agent environments. Moreover, the evaluated scanners fail to detect the constructed skills, motivating behavioral auditing of reusable skills as agent policy artifacts.
Jia-Rui Li, Jiahao Chen, Chunyi Zhou et al.· 0 citations
This work forms this problem as backdoor generalization under training--inference trigger shift and introduces Lilith, a black-box anchor-to-family framework that achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap.
Unsafe Semantic Distillation is proposed, which aligns adversarial perturbations with distributional representations of unsafe content rather than prompt-specific instances, and achieves 84% attack success rates, outperforming existing methods and exposing fundamental vulnerabilities in current multimodal safety architectures.
Shuo Shi, Ruiping Yin, Naen Xu et al.· Proceedings of the 32nd ACM...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.