Skip to content
Preprint

Plausible Deniability Guarantees for Whistleblowers

Jul 2026 · 0 citations · 38 references
Computer Science Mathematics

TL;DR

This work formalizes protection against a strong-adversary threat model as per-report $(0, \delta)-differential privacy on the transcript of audit selections, and proves that a natural approach can never outperform uniform random auditing by more than $\delta$ at any horizon.

Abstract

Whistleblowers are a key safeguard against organizational wrongdoing, but the threat of retaliation deters reporting. Existing whistleblower-protection proposals lack formal privacy guarantees, and existing differential privacy mechanisms do not directly target the natural threat model -- one in which the audited organization itself observes auditor selection decisions and uses them to identify reporters. We formalize protection against a strong-adversary threat model as per-report $(0, \delta)$-differential privacy on the transcript of audit selections. Within this framework we prove that a natural approach -- randomized response applied at the selection step -- can never outperform uniform random auditing by more than $\delta$ at any horizon. We then give a generic mechanism that reduces private auditing to private continual counting: any $(0, \delta)$-DP continual counter plugs in by post-processing, and the audit transcript inherits the same per-report guarantee. Instantiating the reduction with a recent work in continual counting yields per-report $(0, \delta)$-DP with noise scaling as $O(\sqrt{\log T})$ across a horizon of $T$ audit decisions. A utility theorem shows that the selection error vanishes whenever the noisy report gap between the most-reported organization and the runner-up grows faster than $\sqrt{\log T}$. Simulations show a substantial improvement over randomized response.

View source

Similar papers

Preprint Aug 2026

Manipulation-Proof Oblivious Audits against Deceptive Model Providers

A novel audit protocol designed to significantly increase the post-audit detectability of manipulations by enabling the auditor to query the model in an oblivious manner and providing theoretical guarantees showing that, under this protocol, a provider attempting to hide unfairness must falsify a significantly larger number of responses.

Augustin Godinot, Sofiane Azogagh, Julien Ferry et al. · 0 citations
Jul 2026

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited. We introduce a lightweight black-box auditing framework that injects four silent failure profiles across 12 production-adjacent tool stubs and classifies agent responses into three mutually exclusive behavioral classes: Honest Surrender (HSR), Fabrication (FAR), and Unfaithful Safety Refusal (USR). Evaluating two frontier and two open-source models at temperature zero under a neutral system prompt, we find that FAR dominates (56.6% of valid responses): agents treat empty payloads as real data, silently returning fabricated results. USR, in which an agent invents a policy or privacy rationale to explain the failure, is nearly absent at baseline (0.25%, one instance across 396 valid trajectories). Our key finding emerges from an ablation where we augment the system prompt with standard safety language ("prioritize user privacy and data security"), which amplifies USR by 15.6x (from 0.25% to 3.95%; 95% CI on ablation rate: 2.2%-6.4%; Fisher's exact test, p<0.001). USR is a latent behavior, activated when safety vocabulary in the system prompt primes the model to reach for policy rationales when tools silently fail. Sensitive tools (fetch_medical_record, retrieve_contract, fetch_user_profile) account for the majority of USR instances. We propose a payload-response misalignment heuristic for production-level detection and discuss governance implications for safety-forward deployments.

Aarushi Singh · 0 citations
Preprint Aug 2026

Beyond Best Response: Quantal Stackelberg Deception as Insurance Against Attacker Misspecification

Stackelberg Security Games (SSG) assume that an attacker observes the defender's strategy and chooses the target that maximizes their expected utility perfectly. In most realistic applications this is not plausible, and in the case of cyber deception (e.g., using decoys) the purpose of the game is to induce uncertainty and mistakes. Quantal response is a common way to represent noise and mistakes in decision-making; here it replaces perfect best-response with a logit choice with rationality parameter $\lambda$ and results in a generalized Quantal Stackelberg Equilibrium (QSE), which recovers the classical solution exactly as $\lambda \rightarrow \infty$. We conduct a deeper analysis of how QSE can function as a generalized form of insurance against a variety of forms of model specification error/uncertainty; our analysis shows that QSE provides a practical way to address the important role of tie-breaking rules and model uncertainty in SSG from both a theoretical and practical perspective. We conduct an empirical evaluation in a cybersecurity case study with two networks and real vulnerabilities drawn from CVE and scored using the Common Vulnerability Scoring System (CVSS). QSE beats Stackelberg in realized defender utility spanning 144 scenarios with specification errors and 25 parameter configurations, with gains of 46\% to 175\% showing a substantial advantage in a wide variety of realistic cases.

Asif Rahman, Md Abu Sayed, Ahmed Ann Noor Ryen et al. · 0 citations
Review Aug 2026

Repeated-Game Security for Restaking-Based Verifiable Inference

A deployable mechanism combining history-dependent challenges, reputation-weighted slashing, and stake vesting is proposed, which restores infinite-horizon subgame-perfect incentive compatibility against stationary mixed-strategy deviations above an explicit discount-factor threshold without per-query cryptographic verification.

Zhenhang Shang, Yingzhe Yu, Kani Chen · 0 citations
#artificial intelligence Preprint Aug 2026

GuardianAgent: Policy-Conditioned Risk-Adaptive Anonymization with Verified Adversarial Escalation

GuardianAgent, a policy-conditioned anonymization framework that couples structured risk assessment with verified adaptive rewriting, achieves the strongest privacy-utility trade-off among published baselines and is the only method to reach more than 0.90 privacy in all three domains, remaining robust under a backbone switch.

Ruiyi Yang, Gayathri Lihinikaduarachchi, Rahat Masood et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.