Skip to content
Preprint

Bessel-Debiased Pseudo-Marginal MCMC for Generalised Bayesian Inference

Aug 2026 · 0 citations
Mathematics

TL;DR

Sign-Corrected Bessel Debiasing (SCBD), a signed pseudo-marginal method based on independent block estimates of the loss, is introduced and its ordinary-MC and independently randomised quasi-Monte Carlo (RQMC) implementations are studied.

Abstract

Generalised Bayesian inference uses weights of the form $\exp\{-\beta_n\ell_n(\theta)\}$ even when the loss is only estimated. Exponentiating an unbiased loss estimate changes the target, and when $\beta_n\asymp n$ an ordinary Monte Carlo loss estimate with variance of order $M^{-1}$ requires a per-proposal budget $M$ of order $n^2$ to keep the leading log-weight variance bounded. We introduce Sign-Corrected Bessel Debiasing (SCBD), a signed pseudo-marginal method based on independent block estimates of the loss, and study its ordinary-MC and independently randomised quasi-Monte Carlo (RQMC) implementations. Under an i.i.d. Gaussian block model, a Bessel factor constructed from the block sample variance exactly removes the Gaussian exponential inflation despite the variance being unknown. For general non-Gaussian finite blocks, the method targets a posterior differing from the intended posterior by a parameter-dependent multiplicative factor. Under regularity conditions, the uncorrected and corrected ordinary-MC targets have total-variation errors of orders $\beta_n^2/M_n$ and $\beta_n^3/M_n^2$. If an RQMC block estimator has variance $\mathcal O\{B^{-\alpha}(\log B)^{d-1}\}$, the corresponding errors are of orders $\delta^{\mathrm{RQ}}_{n,M_n}$ and $(\delta^{\mathrm{RQ}}_{n,M_n})^{3/2}$, where $\delta^{\mathrm{RQ}}_{n,M_n}=\beta_n^2M_n^{-\alpha}{\log(2+M_n)}^{d-1}$. The same variance rate gives a sufficient budget of order $n^{2/\alpha}$, up to logarithmic factors, for bounded leading log-weight variance when $\beta_n\asymp n$. The numerical examples show that favourable RQMC representations can inherit this budget scaling and that variance reduction and Bessel correction are complementary. Compared to existing exact corrections, Bessel debiasing is essentially"for free". It is generic, easy to code and supported by theory.

View source

Similar papers

#machine learning Preprint Sep 2026

Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling

We study a variant of the Thompson Sampling (TS) algorithm, called $\alpha$-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We formalize the idea of variance inflation by introducing $\alpha$-TS that uses a fractional or $\alpha$-posterior instead of the standard posterior. Our main contribution is to identify general regularity conditions on the prior and reward distributions that enable a regret analysis of $\alpha$-TS without assuming any tractable approximation of the posterior distribution, unlike previous works. For a specific choice of $\alpha \propto d^{-1}$, our general regret bound yields the best known regret bound of $O(d^{3/2}\sqrt{T}\log T)$ for both the exponential and sub-Gaussian families of reward distributions. We further provide an $\alpha$-dependent lower bound showing that the regret constant depends on the product $\alpha d$, and that when $\alpha \propto d^{-1}$ the regret scales as $\Omega(d^{3/2}\sqrt{T})$, explaining the origin of the $d^{3/2}$ factor in the upper bound. Our proof technique adapts and combines recent advancements in the analysis of linear bandit problems with first- and second-order posterior concentration theory from the Bayesian statistics literature.

Prateek Jaiswal, D. Pati, A. Bhattacharya et al. · 0 citations
Preprint Aug 2026

Information-Computation Inversion in Pseudo-Marginal MCMC

An inverse-weight acceptance bound is derived that yields a lower bound on functional mean-squared error from high retained-weight mass in discrepant states and distinguishes statistical information, retained-state functional risk, and auxiliary-computation allocation in pseudo-marginal inference.

Zihan Xu · 0 citations
Preprint Aug 2026

The Optimal Discounting Parameter of the Power Prior under Predictive Log-Loss

The power prior of Ibrahim and Chen incorporates historical data into a Bayesian analysis by raising the historical likelihood to a power $a_0 \in [0, 1]$. The choice of the exponent has remained an open question. This paper gives a closed-form answer under the predictive log-loss. For a model with $d$ parameters, a historical sample of size $N_0$, and average Kullback--Leibler divergence $\bar{D}_0$ between the historical and current data-generating distributions, the optimal exponent is $a_0^{*} = d/(2 N_0 \bar{D}_0 + d)$. Equivalently, the optimally borrowed effective sample size obeys the harmonic law $1/E^{*} = 1/N_0 + 2\bar{D}_0/d$: compatible data are pooled in full, and any difference caps the borrowed information at $d/(2\bar{D}_0)$ observations. The result is exact for multinomial data and extends to smooth parametric families. The law benchmarks adaptive borrowing, explains the reported degeneracy of the normalized power prior, and shows that neither subsetting the data nor decaying the exponent improves on the correctly discounted constant.

Yuriy A. Reznik · 0 citations
Preprint Jul 2026

Amortized Inference for Sampling Distributions Where the Bootstrap Fails

A neural network is trained on simulated datasets drawn from a prior over a distribution family, using single independent draws of the root T_n - T(F) scored by the pinball loss, a proper scoring rule whose population minimizer is the posterior-predictive law of the root.

Akash Deep · 0 citations
Preprint Jul 2026

Local permutation tests for conditional independence: an adaptive binning perspective

In this work, we study the problem of testing conditional independence between random variables $X$ and $Y$ given a confounder $Z$. The local permutation test (LPT) offers a principled approach to this problem by partitioning the $Z$-space into pre-specified bins, and permuting the $X$ and $Y$ data within each bin, to assess the significance of an observed test statistic. However, when the partitions are pre-fixed, the resulting partition can be poorly balanced, as some bins may contain most of the samples while others contain only a few. This motivates the use of data-adaptive binning strategies, such as equisized bins with a fixed (typically small) number of points. We study this natural and practically important extension of LPT, providing finite-sample bounds on the Type I error for an arbitrary test statistic, providing stronger validity results than previously known. We also show that LPT attains power comparable to the oracle likelihood ratio tests derived from the Neyman-Pearson lemma. Within a linear confounder model class, we further analyze the effect of bin size and demonstrate that constant bin sizes can match the performance of partitions with growing bin-size. These results, further supported by extensive numerical simulations, position the proposed data-adaptive strategy as both practically implementable and statistically efficient.

David Chen, Rohan Hore, R. Barber · 0 citations
Preprint Aug 2026

Breakdown Reliability for Saturated Fixed-Effect Inference

Fixed-effect saturation does not itself distort conventional inference, but classical measurement error does. Under a local noise drift $\sigma_\nu^2=c^2/n$, the FE-OLS $t$-statistic converges to a non-central normal; saturation contributes a common $\sqrt{1-\rho}$ scaling rather than preferentially destroying signal or noise. Inverting the size distortion gives a Stock--Yogo-style critical value. Self-consistency of the within-reliability-corrected pilot yields a breakdown reliability $\lambda^{\dagger}=|t|/(|t|+\eta^{\dagger})$ --- the minimum within reliability at which conventional inference retains nominal size within the chosen tolerance --- computable from the reported $t$-statistic alone and algebraically $\rho$-free conditional on it; $\eta^{\dagger}\approx0.65$ at $5\%$ size and a 5-point tolerance. Replacing $|t|$ by $|t|+z_{1-\gamma_\beta}$ gives a certified breakdown reliability; with a lower-reliability bound whose coverage error is $\gamma_\lambda$, false certification is at most $\gamma_\beta+\gamma_\lambda$. Under a checkable projection-compatibility condition, a cluster-level score CLT and consistency of the Arellano variance estimator in the many-fixed-effect regime justify applying the same map to the reported cluster-robust $t$-statistic; clustering can reverse a verdict. In a saturated democracy--growth panel, aggregate V-Dem polyarchy is certified at $\gamma_\beta=0.05$, conditional on the supplied measurement model, while its judicial-constraints sub-index is flagged under i.i.d.\ and clustered standard errors. In a twin-pair wage design, the specification is flagged under both independent and correlated reporting-error models, although implied coverage of the nominal-$95\%$ interval ranges from $8\%$ to $68\%$. The diagnostic covers classical error in a continuous regressor, not binary-treatment misclassification.

S. Halkiewicz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.