Results support auditable, reversible control over a trained parameter path, while showing that useful transfer remains distribution dependent.
Abstract
Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by the forward pass? Policy-Masked Private Experts freezes a pretrained sparse Mixture-of-Experts (MoE) model, trains a disjoint expert branch, and selects the public or private pool before top-k routing. The resulting claim is narrow but testable: under the declared trusted computing base (TCB), an unauthorized request executes no private expert. It does not imply that the public model lacks the same semantic capability. We test this separation between execution control and task utility in Qwen3-30B-A3B and DeepSeek-V2-Lite. Three Qwen BF16 seeds update all 32 private experts while the public fingerprint remains unchanged. Across 64 adversarial scenarios and 96 deny/fail-closed events, unauthorized private execution is zero; independent hooks exactly match 11,616 routed private rows and allow-deny-allow recovery is exact. On two prospectively frozen Qwen benchmarks, the private branch improves exact tool use by 5.0 percentage points (pp) (five versus zero discordances; one-sided Holm p = 0.03125, corresponding two-sided exact p = 0.0625) and 21.3 pp (percentile-bootstrap 95% CI [13.3, 29.3], Holm p = 0.000031). Three arm-blinded model evaluators retain a positive external effect of 18.7 pp (95% CI [9.3, 28.0]). A parameter-matched Lora has similar external utility, but a post-hoc request gate leaves 1,225 adapter calls under deny; the disjoint expert branch leaves none. DeepSeek reproduces the route invariant and gains 27.0 pp. A valid sealed evaluation is near-neutral. These results support auditable, reversible control over a trained parameter path, while showing that useful transfer remains distribution dependent.
TriShield is presented, a three-layer deterministic defense that completely prevents NeuroImprint-style reconstruction with zero model utility loss and no additional communication rounds, and it is proved theoretically that after Layers 2 and 3, the mutual information between the uploaded gradient and any individual training sample is zero.
Gecko is presented, designed to limit this additional risk while retaining a compact encrypted predictor, and formalizes ideal independence and information-preservation conditions as design guidance, then separately evaluate component-reuse extraction attacks.
Cheng'an Wei, Kai Chen, Yue Zhao et al.· 0 citations
Public ledgers increasingly authorize state transitions using prior transactions, finalized state, timing, and ordering rather than only a public key, message, and portable signature. We introduce ledger authenticators and $\LAEUF$, an unforgeability experiment for reactive authorization protocols whose public judgment algorithm reads a finalized transcript. The model separates authentication safety from ledger liveness and captures canonical transition freshness, adaptive corruption, exposure before inclusion, censorship, and adversarial ordering. We identify two conditional resource boundaries. An authenticator satisfying our single event conditions yields a contextual one-time signature. Within our rebindable reveal class, safety requires computational post-disclosure non-admissibility. When precursor admission uses only public computation and ledger scheduling, this condition is enforced by closing the evidence eligible to use a disclosed credential. If newly constructed evidence remains admissible after disclosure, censoring the honest reveal gives a forgery. We then define a joint ledger and quantum random oracle execution model in which quantum state persists across classical finalization cuts and oracle evaluations made through the ledger are charged. For a closed finalized target set of size at most $K$, we prove the bound $3\beta_{\mathsf{cut}}^2+3c_{\mathsf{co}}KQ^2/2^\lambda+6\ell/2^\lambda$, where $\beta_{\mathsf{cut}}$ accounts for fresh openings already present at the cut. A commit, close, reveal authenticator instantiates the framework and obtains a multi-user lifetime QROM bound.
This work expands the understanding of attacks in this setting by investigating a broader class of functionalities, namely: Private Set Union, PSU-Cardinality, and Meta’s multi-key private matching (MKPM) functionality, and investigates possible mitigations for deploying such systems.
Andrea Raguso, Francesca Falzon, Tianxin Tang et al.· Proceedings on Privacy Enhan...· 0 citations
A privacy-preserving zk-SNARK-based audit framework that searches for probes designed in the spirit of adversarial examples to amplify logit drift between an approved model and a modified deployment and demonstrates that token-based probes consistently deliver the strongest mean sensitivity across models and GPU platforms, although operating in a black-box setting.
Cameron Wilding, Mina Shaker, Fatemeh Ganji· 0 citations
Existing Capture-the-Flag (CTF) platforms trust a single organizer, offer limited auditability, and are vulnerable to infrastructure-level manipulation. We propose zk–MPSFV, a zk-SNARK-based, multi-phase sub-flag verification scheme that replaces centralized scoring with an on-chain, zero-knowledge, publicly verifiable scoreboard. Challenges are decomposed into sub-challenges arranged as a directed acyclic graph (DAG): a team unlocks the next step only after proving completion of all parent nodes. Sub-flags and decryption keys are jointly generated by n organizers and released via an off-chain ($t, n$) Shamir–BLS threshold signature produced through multi-party computation (MPC), preventing any single organizer from leaking or altering keys. Teams submit zk-PLONK proofs that the contract verifies, timestamps, and records immutably. Under standard assumptions (collision-resistant hashing, SNARK soundness/zero-knowledge, IND-CCA2 ECIES, and at least t honest organizers), we prove that zk–MPSFV achieves the stated security goals, including DAG-gated progress, anti-replay, and threshold-robust organizer security, while out-of-band flag sharing remains out of scope. On a three-organizer testbed with 30 simulated teams, setup costs 0.45 ms per sub-flag, proof generation averages 5.34 s on an 8-core system, and on-chain verification costs $\approx$ 170kL2 gas on zkSync Era with a median fee of 1.33 $\times 10^{-6}$ ETH (about ${\$}$0.0046 at ${\$}$3,435/ETH). Stress replays sustain $\approx$ 7 proof transactions/s up to 5000 proofs; extrapolating to 50,000 proofs (1000 teams $\times$ 50 submissions) yields $\approx$ 0.0665 ETH (about ${\$}$200–${\$}$228) and $\approx$ 2 hours of settlement time. Overall, zk-MPSFV is practical for small- to mid-scale, audit-ready progression CTFs.
S. Khanji, Behzad Abdolmaleki, John A. Clark et al.· IEEE Computer Security Found...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.