Skip to content

Tamed Stochastic Gradient Hamiltonian Monte Carlo

Jul 2026 · arXiv.org · Vol abs/2607.14862 · 0 citations
Computer Science Mathematics

TL;DR

A novel tamed stochastic gradient Hamiltonian Monte Carlo algorithm for sampling and stochastic optimization problems with superlinearly growing stochastic gradients is proposed, which achieves lower root mean square error and expected excess risk across a range of tasks.

Abstract

In this paper, we propose a novel tamed stochastic gradient Hamiltonian Monte Carlo (tSGHMC) algorithm for sampling and stochastic optimization problems with superlinearly growing stochastic gradients. Under a certain continuity in average condition and a strong convexity condition, we establish a non-asymptotic error bound in Wasserstein-2 distance for tSGHMC with the rate of convergence equal to $1/4$. Then, we derive an upper estimate for the associated expected excess risk, which provides a theoretical guarantee for the performance of tSGHMC. To illustrate the effectiveness of the proposed algorithm, we apply tSGHMC to practical examples, including a newsvendor problem and a Conditional Value-at-Risk minimization problem, using synthetic and real-world datasets. Numerical results support our theoretical findings. Furthermore, we compare tSGHMC with its first-order counterpart, namely, the tamed unadjusted stochastic Langevin algorithm. Simulation results demonstrate that tSGHMC achieves lower root mean square error and expected excess risk across a range of tasks.

View source

Similar papers

Open access Aug 2026

Unbiased kinetic Langevin Monte Carlo with inexact gradients

Theoretical analysis demonstrates that the proposed estimator is unbiased, attains finite variance, and satisfies a central limit theorem, and the results demonstrate that in large-scale applications, the unbiased algorithm can be 2–3 orders of magnitude more efficient than the “gold-standard” randomized Hamiltonian Monte Carlo.

Neil K. Chada, B. Leimkuhler, Daniel Paulin et al. · 0 citations
Preprint Aug 2026

Zeroth-Order Langevin Monte Carlo via SPSA under Noisy Function Measurements

In sampling problems, gradient-based schemes such as Langevin Monte Carlo (LMC) mix faster than non-gradient-based methods, but their applicability is limited by access to the gradient of the target log-density. In practice, gradients are often unavailable and function evaluations are noisy, e.g., stochastic simulators or black-box simulators, so we propose LMC-SPSA with noise, which approximates the gradient of the target log-density using two noisy function evaluations per iteration. We prove, under noisy gradient estimates, that LMC-SPSA converges in distribution by proving the convergence in Wasserstein distance. Furthermore, we construct a diminishing step-size schedule that still drives the Wasserstein error bound to convergence, extending convergence guarantees beyond the constant-step setting. Further, we sharpen the dominant dimension dependence of the Wasserstein error from $O(p^4)$ to $O(p^2)$ (with $p$ denoting the dimension), and support this analysis with numerical results. We show that LMC-SPSA achieves $W_2$-accuracy $\varepsilon$ with total noisy-oracle complexity of $O(p/\varepsilon^2+\delta^2p^3/\varepsilon^3)$, where $\delta$ is the paired-noise level. This improves the noise-dependent accuracy scaling relative to the ZO-LMC method of Roy et al. We further establish asymptotically vanishing Wasserstein error as the number of iterations $\to\infty$ under diminishing step-size and perturbation sequences and derive an explicit convergence rate for a balanced schedule under noisy zeroth-order feedback. Empirical experiments are conducted to verify the performance of LMC-SPSA with noise. We provide an oracle-budget-matched comparison with the ZO-LMC method, showing smaller empirical sampling errors under the same function-evaluation budget.

Hongbo Li, J. Spall · 0 citations
Preprint Aug 2026

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.

Iosif Lytras, Nikolaos Makras, S. Sabanis · 0 citations
Preprint Jul 2026

Sinkhorn Hamiltonian Monte Carlo for Entropic Optimal Transport Generalized Bayes

Bayesian posterior sampling is a ubiquitous paradigm for problems where a point estimate of parameters is not sufficient, such as risk analysis and uncertainty quantification. However, likelihoods may be misspecified, intractable, computationally expensive, or not representative of the discrepancy of interest. Generalized Bayes extends likelihood-based posterior updates by using other losses. Sinkhorn divergences have appealing geometric properties: they compare empirical measures directly and yield smooth gradients thanks to entropic regularization. In this work, we introduce Sinkhorn divergences as Generalized Bayes losses for Hamiltonian Monte Carlo (HMC) and No-U-Turn Sampler (NUTS). We also propose heuristics to set hyperparameters that affect the stability and calibration quality, such as the number of Sinkhorn iterations, the entropic regularization strength, and the marginal relaxation penalty. In regimes where the forward model relies on a stochastic simulator, we combine HMC/NUTS with a common-random-numbers strategy to obtain a deterministic surrogate objective that preserves gradients and Hamiltonian dynamics. We study both mass-preserving balanced and relaxed unbalanced settings. We evaluate our method empirically on (1) a simple Gaussian model as a sanity check; (2) a distribution supported on a noisy spiral manifold where a likelihood-based approach is a poor fit; (3) a Gaussian pulse model with misalignment due to errors-in-variables, emphasizing robustness to misspecification; and (4) CIFAR-10 image patch alignment under perturbations, highlighting differences between balanced and unbalanced regimes.

Guilhem Nespoulous, F. Bertrand, Myriam Maumy et al. · 0 citations
Preprint Aug 2026

Stochastic gradient descent with initial regularization

A variant of stochastic gradient descent with initial regularization with initial regularization is analyzed and dimension-free upper bounds on its expected excess risk for the squared loss are derived.

Nabil Kahalé · 0 citations
Preprint Jul 2026

Stochastic Quantization as Optimal Control

Stochastic quantization defines a Euclidean quantum field theory as the equilibrium of a fictitious-time Langevin dynamics, which reaches the Gibbs measure asymptotically and is formulated as a finite-time stochastic optimal control problem.

L. Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.