Skip to content

PAC-Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors

Jul 2026 · arXiv.org · Vol abs/2607.18422 · 0 citations · 16 references
Computer Science Mathematics

TL;DR

It is shown that PAC--Bayesian analysis should be performed on the quotient predictor space: pushing a prior and posterior to the quotient preserves the empirical and population Gibbs risks while removing the nonnegative KL contribution caused solely by how the two distributions differ among parameterizations of the same predictor.

Abstract

Overparameterized models often have continuous parameter symmetries, so different parameters define the same predictor. We show that PAC--Bayesian analysis should be performed on the quotient predictor space: pushing a prior and posterior to the quotient preserves the empirical and population Gibbs risks while removing the nonnegative KL contribution caused solely by how the two distributions differ among parameterizations of the same predictor. Quotienting alone does not determine which prior to use. We construct a canonical choice of one parameterization for each predictor and account for the geometric volume of its equivalent parameterizations. This transforms a neutral reference prior into a data-independent prior that reflects the model's implicit bias. It approximates the ideal but inadmissible posterior-matched prior, which would minimize the KL term by depending on the training data. The resulting certificate is tighter exactly when this geometry-induced prior has smaller KL divergence from the learned quotient posterior than the neutral prior. We test this prediction in Fourier regression with a Hadamard parameterization and in Query-Key attention, using ordinary SGD without an explicit regularizer. The implicit-bias prior reduces the mean quotient-space KL by \(40.69\%\) and the mean PAC--Bayes certificate by \(21.40\%\) in the Fourier-Hadamard experiment. The smaller, prior-scale-dependent improvement in Query-Key attention confirms the predicted conditional nature of the effect.

View source

Similar papers

Preprint Sep 2026

On Prior-to-Posterior Stability in the Wasserstein Metric for Bayesian Inverse Problems

Priors in Bayesian inverse problems are often approximated through discretization, hyperparameter estimation, or generative modeling. Understanding how prior approximation errors propagate to the posterior and subsequent predictions is therefore important. In this work, we study the stability of the prior-to-posterior map where both prior and posterior perturbations are measured in the same Wasserstein metric $W_p$, $p\geq1$. We identify verifiable conditions on likelihood regularity and admissible prior classes that ensure uniform, H\"older, and Lipschitz stability. For bounded likelihoods that are uniformly continuous on bounded sets, uniform stability holds over prior classes with uniformly integrable $p$-th moments and a common positive evidence lower bound. With global H\"older regularity of the likelihood and uniform bounds on higher prior moments, a coupling argument leads to a H\"older estimate with a sharp exponent. For $p>1$, a Lipschitz likelihood need not give Lipschitz stability, even for priors with bounded support. We establish Lipschitz stability through an interpolation argument under uniform Poincar\'e bounds and a globally Lipschitz potential with uniformly bounded essential oscillation under the priors. For Gaussian priors with additive Gaussian noise and bounded Lipschitz forward models, these estimates give posterior $W_2$ bounds even for mutually singular prior perturbations. Numerical experiments for a Darcy inverse problem illustrate the predicted H\"older and Lipschitz rates and the resulting control of errors in the posterior mean and standard deviation of a Lipschitz quantity of interest.

Unknown authors · 0 citations
Preprint Aug 2026

PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition

PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation. However, predictive risk depends only on the predictive behavior induced by a hypothesis, not on the particular internal realization that implements that behavior. In over-parameterized systems, many distinct configurations induce identical predictive behavior, yet the classical PAC-Bayes KL divergence does not distinguish uncertainty over predictive behavior from variation among behaviorally equivalent realizations. We show that this distinction induces an exact structural decomposition of classical PAC-Bayes complexity. We formalize behavioral equivalence through a measurable behavior map and use measure disintegration to decompose probability measures on the configuration space into a distribution over predictive behaviors and conditional distributions over behavioral fibers. This yields an exact decomposition of the classical PAC-Bayes KL divergence into a behavior-selection term and a realization-level term given by an expected conditional KL within fibers. We define Z-information as the negative of this realization-level contribution: the exact gap between the KL divergence and the complexity of uncertainty over predictive behavior alone. We further show that the behavior-selection term admits an exact variational characterization: it is the minimum KL divergence among all posteriors inducing the same distribution over predictive behaviors, attained by a canonical fiber-symmetrized representative. Finally, we show that symmetry, behavior-preserving directions, fiber geometry, and invariance under fiber-preserving perturbations arise naturally from the same behavior-map structure. Together, these results identify predictive behavior as the natural object of PAC-Bayes complexity.

Vasant Honavar, Satish Kumar Keshri, Neil Ashtekar et al. · 0 citations
Preprint Aug 2026

Duality and Error for Predictively Oriented Inference

This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space.

Aurya Javeed, D. Kouri, Teresa Portone et al. · 0 citations
Preprint Aug 2026

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

This work presents the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynomial neural networks.

Grégoire Sergeant-Perthuis, E. Tsigaridas, Jules Tsukahara Cqsb et al. · 0 citations
Preprint Aug 2026

Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families

We study posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. We embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link condition on the Fisher information and suitable local regularity assumptions, we show that smoothness-matching priors achieve minimax-optimal posterior contraction rates in any Hilbert scale norm up to the regularity of the ground truth. Our analysis builds on the novel approach to posterior contraction based on the Wasserstein distance recently introduced by Dolera et al. (2024). It combines refined Laplace-type estimates for infinite-dimensional integrals associated to the posterior kernels with a mixed-geometry estimate controlling their stability under fluctuations in the data, itself resting on a tailored Poincar\'e inequality for posterior distributions conditioned on neighbourhoods of the truth. We apply the general theory to density estimation with a logistic parametrisation, Poisson intensity estimation with an exponential link, and the Gaussian white-noise model, yielding minimax contraction rates in Sobolev norms across all three settings. In particular, these yield optimal recovery of density score functions and derivatives of Poisson intensities.

Emanuele Dolera, Stefano Favaro, M. Giordano · 0 citations
#machine learning Preprint Sep 2026

Generalized Score Matching for Parameter Estimation on Convex Domains

Maximum likelihood (ML) estimation is a principled and statistically efficient approach for learning probabilistic models. However, for unnormalized models, ML estimation requires evaluating the partition function and differentiating through it, which may not always be tractable. Score matching provides a practically viable alternative that circumvents this obstacle by fitting the score in a way that eliminates dependence on the normalizing constant. We derive the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and show how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework. We show that the resulting objective is a {\it proper local scoring rule} of second-order, which provides the theoretical guarantee that the true density is recovered when the objective is minimized. Furthermore, for a model belonging to the exponential family, we establish convexity of the objective together with consistency of the finite-sample estimator under standard regularity conditions. Our derivation sheds new light on the scope and applicability of generalized score matching in various problem settings. We compare generalized score matching-based estimators on constrained domains, where the partition function is analytically intractable. We provide experimental results on parameter estimation for model densities belonging to the exponential family defined over convex subsets of $\mathbb{R}^{d}$, and a generative modeling use-case to demonstrate broader applicability of the proposed generalized score matching framework.

Nishanth Shetty, Saisuchith Mahajan, C. Seelamantula · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.