Skip to content
Preprint

Exact Algebraic Computation of Learning Coefficients for Two-Dimensional Singular Models

Aug 2026 · 0 citations · 62 references
Computer Science Mathematics

TL;DR

This work presents the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynomial neural networks.

Abstract

Classical information criteria such as the Bayesian Information Criterion (BIC) rely on regularity assumptions that break down for singular models, leading to incorrect model selection in settings such as deep learning. The Widely Applicable Bayesian Information Criterion (WBIC) relies on local learning coefficients $\lambda$, which in the analytic case coincides with local Real Log Canonical Thresholds (RLCT) of the Kullback-Leibler divergence of the model, to capture correct marginal likelihood asymptotics. Exact computation of the learning coefficients has been limited to special cases, and only sampling-based estimation methods are generally applicable. We present the first deterministic algorithm that computes local RLCTs exactly for any two-dimensional model whose Kullback-Leibler distance is contact equivalent to a polynomial, derive a bound on its complexity, and demonstrate its effectiveness for a broad class of models, with applications including polynomial neural networks. Beyond providing ground truth to calibrate sampling-based estimators, exact computation reveals algebraic structure in learning coefficients that sampling cannot and out-speeds it in the shallow regime.

View source

Similar papers

Preprint Sep 2026

On the sample complexity of the active subspace method

Active subspaces identify low-dimensional linear structure in high-dimensional parameter-to-output maps by estimating the dominant eigenspace of a gradient covariance operator. In practice this covariance is replaced by a Monte Carlo estimator built from a limited number of gradient evaluations. Classical analyses based on controlling the covariance error in operator norm lead to sample-complexity estimates that can be substantially more pessimistic than the sampling rules commonly used in computations. This paper studies the empirical active subspace method directly in the projection-error metric relevant for ridge approximation. We derive non-asymptotic quasi-optimality bounds governed by a regularized inverse Christoffel function associated with the gradient field. Under a bounded-gradient assumption, the resulting estimates already improve the sample-complexity estimates obtained from operator-norm covariance bounds. We then show that additional smoothness of the gradient map, expressed through membership in a reproducing kernel Hilbert space, yields sharper coherence estimates and motivates tractable importance sampling from kernel diagonal measures. Furthermore, the same smoothness assumption yields a priori decay bounds for the population active subspace tail energy, which can be combined with our finite-sample estimate to prescribe rank, regularization scale, and sample size, allowing to fully characterize the a priori sample complexity. The abstract assumptions are verified for lognormal Gaussian and affine uniform parametric elliptic PDEs using weighted summability of Hermite and Legendre series expansions.

Fabio Nobile, Matteo Raviola, R. Tempone · 0 citations
Jul 2026

PAC-Bayes Bounds on Quotient Parameter Spaces: Geometry-induced Implicit-Bias Priors

It is shown that PAC--Bayesian analysis should be performed on the quotient predictor space: pushing a prior and posterior to the quotient preserves the empirical and population Gibbs risks while removing the nonnegative KL contribution caused solely by how the two distributions differ among parameterizations of the same predictor.

Nicola Aladrah, Fabio Anselmi · 0 citations
#machine learning Preprint Sep 2026

Generalized Score Matching for Parameter Estimation on Convex Domains

Maximum likelihood (ML) estimation is a principled and statistically efficient approach for learning probabilistic models. However, for unnormalized models, ML estimation requires evaluating the partition function and differentiating through it, which may not always be tractable. Score matching provides a practically viable alternative that circumvents this obstacle by fitting the score in a way that eliminates dependence on the normalizing constant. We derive the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and show how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework. We show that the resulting objective is a {\it proper local scoring rule} of second-order, which provides the theoretical guarantee that the true density is recovered when the objective is minimized. Furthermore, for a model belonging to the exponential family, we establish convexity of the objective together with consistency of the finite-sample estimator under standard regularity conditions. Our derivation sheds new light on the scope and applicability of generalized score matching in various problem settings. We compare generalized score matching-based estimators on constrained domains, where the partition function is analytically intractable. We provide experimental results on parameter estimation for model densities belonging to the exponential family defined over convex subsets of $\mathbb{R}^{d}$, and a generative modeling use-case to demonstrate broader applicability of the proposed generalized score matching framework.

Nishanth Shetty, Saisuchith Mahajan, C. Seelamantula · 0 citations
Preprint Aug 2026

Stochastic gradient descent with initial regularization

A variant of stochastic gradient descent with initial regularization with initial regularization is analyzed and dimension-free upper bounds on its expected excess risk for the squared loss are derived.

Nabil Kahalé · 0 citations
#machine learning Preprint Sep 2026

A computational approach to maximum likelihood thresholds for colored Gaussian graphical models

Gaussian graphical models (GGMs) are essential tools for interpretable structure learning. However, in high-dimensional, small-sample regimes, the available data is often insufficient for the maximum likelihood estimator to exist. Colored Gaussian graphical models (CGGMs) mitigate this limitation by imposing symmetry constraints through graph coloring, which reduces the required sample size. This minimal number of observations needed to guarantee that the estimator exists almost surely is defined as the maximum likelihood threshold (MLT). Here, we address the computation of the MLT for CGGMs by focusing on its geometric formulation: finding the minimum rank of a sample covariance matrix such that its projection lies almost surely within the interior of the cone of sufficient statistics. We establish a unified theoretical framework, extending results from uncolored to colored models and introducing new symbolic algorithms. Furthermore, we present a computational study integrating sampling with topological data analysis (TDA) to investigate the local geometry of the cone of sufficient statistics. Our results demonstrate the potential of TDA to overcome the computational bottlenecks of traditional symbolic algebraic methods, particularly Groebner basis computations, in analyzing the likelihood geometry of CGGMs.

R. Homs, Olga Kuznetsova, Bernadette J. Stolz · 0 citations
#machine learning Preprint Sep 2026

How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL

Accurate posterior prediction need not require accurate approximation of Bayesian updates. We prove that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM. Exact Bayesian mixing and an explicit deterministic radial filter act on the same $K-1$ belief coordinates. As $q\to0^+$, their separation in centered logits in the worst case grows at least linearly in the natural confidence scale $L_K(q)$, while their categorical $D_{\mathrm{KL}}(\mathrm{exact}\|\mathrm{radial})$ vanishes at the same explicit witness. Along stationary HMM trajectories, the expected terminal KL between filtered posteriors also converges to zero at $H(q)=\lceil-\log(q)/c\rceil+1$. Typical blocks without switches drive both filters into a common confidence cone, where softmax curvature suppresses their disagreement; a single Gaussian maximal event controls adaptive noise. A sweep with equally spaced Gaussians over $K\in\{2,4,8\}$ illustrates the opposing trends, and binary controls at long horizons compare saturating and nonsaturating recurrences. The result isolates two missing links between internal update gaps and predictive cost: the contribution of separating states to expected loss and decoder sensitivity. Thus even an unbounded internal update gap does not by itself certify predictive failure. The construction is fixed in $K$ and does not provide a universal criterion for when compression is harmless or characterize when internal gaps must incur task loss.

Qi-Fu Wen, Shuai Liu, Zihan Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.