Skip to content
Preprint

Privacy Without Regret: Differentially Private Inference-Time Alignment

Aug 2026 · 0 citations · 35 references
Computer Science

TL;DR

Private Inference-Time Pessimism (PrivITP) is introduced, which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism, and achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses, cleanly decouples the regularization parameter from the privacy parameter, and attains the skyline up to a noise-inflation term.

Abstract

Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward model, and the absence of any privacy protection for the sensitive human preference data used to train that reward model. We show that a single intervention-adding calibrated noise to reward scores before selection-resolves both. Our first result, Private Best-of-N (PrivBoN), establishes that Gumbel noise at an appropriate scale simultaneously provides $\epsilon$-differential privacy and implements KL-regularized alignment. Whenever the privacy budget exceeds a critical threshold $\epsilon^*$, the privacy-mandated noise is the regret-optimal regularization, and privacy imposes zero additional alignment cost-matching the information-theoretic skyline of Huang et al. (2025). Because $\epsilon^*$ depends on an unknown coverage coefficient, we introduce Private Inference-Time Pessimism (PrivITP), which combines $\chi^2$-regularized rejection sampling with a two-phase Gaussian mechanism. PrivITP achieves ex-post $(\epsilon,\delta)$-DP with a privacy cost independent of the number of responses $n$, cleanly decouples the regularization parameter from the privacy parameter, and attains the skyline up to a noise-inflation term. Experiments across several language models, datasets, and reward models confirm our results: PrivBoN and PrivITP are scaling-monotonic (unlike BoN, which degrades past a critical $n$), and PrivITP matches or outperforms PrivBoN at equivalent privacy levels, with the largest gains in the strong-privacy regime.

View source

Similar papers

Preprint Sep 2026

Privacy Amplification Without Independence: How Far Negative Dependence Carries the Guarantees of Poisson Subsampling

Poisson subsampling is the default sampler in differentially private optimization because its independence makes privacy amplification tractable. Practical systems, however, are moving toward structured participation: random allocation (balls-in-bins), per-epoch allocation, random check-ins, schemes widely believed to be at least as private as Poisson subsampling at the matched rate. We isolate the probabilistic mechanism behind this belief and delimit it exactly, for Gaussian mechanisms up to correlated-noise matrix mechanisms. (1) If the participation indicator vector is negatively associated (NA), then at every integer R\'enyi order $\alpha\ge2$, exactly at all finite parameters, its remove-direction R\'enyi divergence is dominated by that of the marginal-matched independent scheme. For fixed gradient sequences, this extends to the mechanism level whenever the noise strategy's Gram matrix is sign-balanced, an $O(t^2)$-checkable condition. (2) The integer-order restriction is essential. For random allocation with $k=1$, we prove a linear law for the R\'enyi-difference criterion: at large $t$, dominance reverses for every $\alpha<3/2$, including KL divergence, while the crossing order tends to $3/2$ independently of $\sigma$. (3) We also localize the known failure of rate-matched Poisson domination exactly: below $(1-q)^t$, the hockey-stick ordering reverses, so substituting the Poisson pair into composition machinery is unsound. An upper-tail argument yields a finite crossover $\gamma_\star$, connecting this threshold picture to the R\'enyi boundary at $3/2$. Together, these results give a substitution map for privacy accounting: when Poisson-based computations remain sound for structured participation, where they fail, and what sound alternatives cost in deployment.

Xujun Che, Depeng Xu · 0 citations
Open access 2025

Optimal Regret of Bandits under Differential Privacy

This work revisits the regret lower and upper bounds of ϵ -global DP bandits and proves a tighter regret lower bound involving a novel information-theoretic quantity characterising the hardness of ϵ -global DP in stochastic bandits.

Achraf Azize, Yulian Wu, Junya Honda et al. · 0 citations
#small language model Preprint Aug 2026

Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling

It is shown that policies within a bounded $\chi^2$ divergence from the proxy-feasible reference distribution admit an $N$-independent safety-hacking bound, and instantiate this general coverage-control principle with constrained pessimistic sampling.

Akifumi Wachi, Takumi Tanabe, Youhei Akimoto · 0 citations
#artificial intelligence Preprint Aug 2026

Performative Privacy: When Differential Privacy Maximizes Utility

It is shown, through a theoretical study of the dynamics and numerical experiments, that a finite privacy budget can outperform non-private estimation in the long term when the feedback loop between leakage and participation is sufficiently strong.

Uddalak Mukherjee, Edwige Cyffers, Y. Chevaleyre · 0 citations
Preprint Aug 2026

Private Generative Bootstrap via Blocking

With AI systems gaining more access to individuals'information, it is important to protect privacy when reporting statistical answers. Equally important is to privatize the reporting of uncertainty in such answers. To this end, we adopt a Bayesian likelihood-free framework and make simulation from the posterior private. In particular, we propose a new private instantiation of the Bayesian bootstrap using a blocking strategy. Rather than assigning idiosyncratic random weights to each individual, we randomly group individuals and assign a single weight to each group. By concealing individuals'contributions within a group, we fortify differential privacy gates. We harness amortized inference that decouples private learning from posterior sampling. A push-forward map from observation weights to posterior samples is learned privately by adding calibrated noise during training. Subsequent posterior draws require no additional privacy and computation budget. We call the resulting method the Private Generative Bayesian Bootstrap (PGBB). We establish a differential privacy guarantee, analyze convergence to the non-private blocked-bootstrap target, and quantify the discrepancy between the ordinary and blocked Bayesian-bootstrap posteriors. In addition, we derive data-free tuning of the block Dirichlet concentration parameter that restores posterior dispersion asymptotically. We also show a single fit of PGBB can support a family of loss-based decision rules simultaneously without additional privacy cost. In simulations and in applications to U.S. Census returns to schooling and U.S. natality birthweight quantiles, PGBB gives competitive private uncertainty quantification and improves over private Bayesian alternatives that require a specified data-generating model in common settings.

Jinwon Sohn, Veronika Rocková · 0 citations
Preprint Aug 2026

Pre-Disclosure Experiment Menus: Oracle-Relative Risk and Joint Sample--Menu Asymptotics

We study a resolution problem in local asymptotic decision theory: individual risks may admit Gaussian approximations that do not determine their vanishing difference. A finite menu of experiments is installed before context disclosure, although observations may be routed adaptively afterward. A greatest-element Blackwell order collapses adaptive routing to the best installed experiment and reduces the fixed-menu excess to an inverse-information distortion with frontier $A_k$. We develop differentiated, all-prior posterior transfer along a one-dimensional degradation chain and establish $F_{n,k_n}(H_n)=A_{k_n}\{1+o(1)\}$ for every diverging menu sequence with positive frontier and every admissible localization radius, without an additional direct sample-menu restriction. The transfer is exact under Gaussian degradation. Prior-free likelihood-generator conditions imply it for jump generators and are verified for binary attenuation, Poisson thinning, and negative-binomial thinning. If the distortion is uniformly quadratic on an Ahlfors-regular oracle image of dimension $r$, then $A_k\asymp k^{-2/r}$, and the original-scale excess mean squared error is of order $n^{-1}k^{-2/r}$. Calibrated Poisson sensor and radial-qubit measurement menus illustrate the result. A triangular Gaussian counterexample shows why pointwise Gaussian convergence is insufficient.

Xinyu Song · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.