It is proved that the success probability of any bounded adversary asymptotically collapses to zero under a defense strategy combining natural absorption, a randomized scheduler, and lazy verification oracle, which yields a provably sound and computationally efficient defense for safety-critical AI.
Abstract
Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to external compute providers within a decentralized training paradigm, adversaries can craft stealthy, low-frequency triggers to inject malicious behavior while evading standard audits. Traditionally, detecting these attacks requires a full re-computation of the training steps--a prohibitive overhead that directly contradicts the owner's resource constraints. To address this, we investigate the resilience of continuous optimization dynamics under Byzantine perturbations, where adversaries are forced to compete against a continuous influx of honest updates. Under a threat model where an adversary compromises f out of n total trainers, we quantify the minimum auditing overhead required by the model owner to probabilistically bound the attack success rate. We formalize this injection-absorption dynamic as a Discrete-Time Markov Chain (DTMC). Using this framework, we prove that the success probability of any bounded adversary asymptotically collapses to zero under a defense strategy combining natural absorption, a randomized scheduler, and lazy verification oracle. Empirical results demonstrate significant backdoor suppression with zero utility degradation even when invoking the verification oracle on merely 10% of the total training steps. This approach yields a provably sound and computationally efficient defense for safety-critical AI.
BackDFL is presented, a unified benchmark for systematically evaluating DFL under realistic and adaptive backdoor attacks, and demonstrates that both state-of-the-art Byzantine-robust DFL methods and adapted FL backdoor defenses fail under modest malicious participation rates, especially in heterogeneous settings.
M. Bouchiha, Gregory Blanc, Yu-Fei Han· 0 citations
TriShield is presented, a three-layer deterministic defense that completely prevents NeuroImprint-style reconstruction with zero model utility loss and no additional communication rounds, and it is proved theoretically that after Layers 2 and 3, the mutual information between the uploaded gradient and any individual training sample is zero.
The game of coding framework was introduced to extend coding-theoretic recovery beyond its traditional limit, under which the number of honest reports must exceed the number of adversarial or corrupted reports. It does so by exploiting the rational behavior of adversarial participants and their incentive to keep the system live. Existing game-of-coding formulations, however, assume that the adversarial-noise distribution is independent of the realized ground-truth computation. This assumption may be restrictive when an informed adversary can adapt its reports to the value being computed. In this paper, we study the game of coding with input-dependent adversarial noise. We introduce a unified multi-node, multidimensional formulation. For every family of conditional adversarial-noise distributions, we construct an input-independent joint noise distribution, and prove that this reduction exactly preserves the probability of acceptance and the accepted mean-squared estimation error. Consequently, the input-dependent and input-independent models have identical achievable performance regions, and the same equilibrium utilities.
Hanzaleh Akbari Nodehi, M. Maddah-ali· 0 citations
A game-theoretic model of steganographic operations that captures the strategic interaction between a defender and an adversary through calibrated monetary primitives and nonlinear utility mappings is presented, and adversarial advantage can be translated into interpretable monetary risk estimates for assessing whether steganographic defences decrease, amplify, or only marginally affect organisational exposure.
Obinna Omego, Farzana Rahman, Onalo Samuel et al.· PeerJ Computer Science· 0 citations
The results indicate that static, benign-traffic-calibrated thresholds are insufficient for this defense, and that jitter-forgiveness thresholds should instead be calibrated dynamically against local token entropy.
Nikita Kezins· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.