Skip to content
Preprint

Markov Chain Monte Carlo with Diffusion Paths

Jul 2026 · 1 citation
Mathematics

Abstract

Sampling from multimodal distributions is a longstanding challenge for classical local Markov chain Monte Carlo (MCMC) methods. A popular remedy is to introduce a sequence of intermediate distributions that interpolate between the target and a simpler reference. The classical choice, tempering, raises the density to a power, but distorts the relative weights of asymmetric modes and can lead to poor mixing. We instead propose interpolating along the diffusion path, the marginals of a noising diffusion process that carries the target toward a Gaussian. This path preserves the relative weights of the modes and enjoys favorable mixing properties, which we make precise through a spectral-gap analysis of the corresponding ideal transition kernel. Sampling along the path requires its intermediate scores, which can be estimated from the unnormalized target through variational approaches, yielding only an approximate sampler. To remove the resulting bias, we introduce the Metropolis-adjusted diffusion path (MAD-Path) sampler, which corrects the diffusion-path proposal in an augmented path space and leaves the target invariant regardless of the accuracy of the learned score or the discretization error. We further quantify how these two errors affect the acceptance probability, providing guidance for practical tuning. Experiments on a range of Bayesian posteriors show that MAD-Path improves global exploration and mode-weight estimation relative to tempering-based MCMC methods and unadjusted diffusion samplers.

View source

Similar papers

#machine learning Preprint Aug 2026

Exact Global MCMC with Denoising Diffusion

This work shows that diffusion models learned with standard denoising loss can provide effective global MCMC proposals for complex high-dimensional target densities. The method is motivated by the observation that sequentially applying a forward and reverse diffusion process defines a Markov chain with a target stationary distribution for an ideal denoiser trained on samples of the target distribution. This observation can be made exact for any denoiser by applying a Metropolis-Hastings step whose acceptance ratio includes the density of the forward and reverse paths of a discrete time SDE approximation. We therefore propose to train denoising diffusion models on locally convergent MALA samples to learn global MCMC proposals. We call the composition of the global denoiser-based path sampler and a local MALA sampler Denoising Diffusion Monte Carlo (DDMC). Experiments show that DDMC can provide global proposals with high acceptance across a variety of complex target densities. Our results offer preliminary evidence that the established scaling behavior of standard diffusion training transfers directly to exact sampling from high-dimensional unnormalized densities.

Mitch Hill · 0 citations
Review Aug 2026

From Continuous Dynamics to Practical Gradient-Based Samplers

Gradient-based Markov chain Monte Carlo methods are often introduced as a catalog of algorithms: Hamiltonian Monte Carlo (HMC), the Metropolis-adjusted Langevin algorithm (MALA), the No-U-Turn Sampler (NUTS), and several underdamped variants. This presentation obscures the common structure of the methods and, more importantly, the reasons why a sampler that is correct in principle may be ineffective in practice. We develop a unified account, beginning with exact continuous-time dynamics that represent idealized sampling methods and for which Metropolis adjustments are not required. Numerical discretization makes the dynamics computationally feasible but introduces bias. Metropolis adjustment removes the asymptotic bias by converting numerical errors into rejection, leading to HMC, MALA, NUTS, and the Metropolis-adjusted kinetic Langevin algorithm (MAKLA). The second half of the paper presents geometric design choices that determine practical performance, namely, although MAKLA and NUTS have nice theoretical properties, their sampling efficiency may be slow in practice. Importantly, a fixed mass matrix can whiten globally anisotropic targets, often fixing sampling inefficiency in Bayesian posteriors with large data. Whereas hierarchical posteriors introduce their own problem, causing state-dependent variation in the Hessian (e.g., Neal's funnel). We explain how a randomized step size can be used effectively to sample from such a distribution. The resulting paper is both a tutorial on the mechanics of gradient-based sampling and a set of practical recipes to improve sampler performance.

J. Chok · 0 citations
Preprint Aug 2026

Robustness of random-walk Metropolis for steep potentials

In Markov chain Monte Carlo sampling, light-tailed target distributions present something of a poisoned chalice: their light tails offer good confinement, and tend to imply good mixing properties for natural continuous-time dynamics, but the steepness of their tail decay means that they often fall outside of the scope of modern quantitative convergence theory. For usual gradient-based samplers, this reflects a genuine instability issue, whereby Metropolis acceptance rates can degrade badly. In this work, we study the gradient-free random-walk Metropolis sampler, and show that for a wide range of light-tailed targets, the acceptance probability remains stable for reasonable choices of proposal variance, from which effective and favourable mixing time estimates can be deduced. The analysis relies on a simple relationship between the first and second derivatives of the log-density of the target distribution.

Sam Power · 0 citations
Preprint Aug 2026

Diffusion Quasi-Monte Carlo

We study high-dimensional numerical integration with respect to complex target measures using diffusion-based transport maps and randomized quasi-Monte Carlo (RQMC). Score-based diffusion models induce a deterministic probability flow ODE that transports a simple prior to the target, suggesting a principled way to transform low-discrepancy points on the unit cube into informative samples. We construct a cube-to-target map by composing a Gaussian base transformation (the component-wise inverse Gaussian CDF) with an Euler-discretized probability flow ODE. To retain unbiasedness under transport approximation, we formulate integration as importance sampling (IS) on the cube. Our main result provides verifiable conditions under which the resulting IS integrand satisfies the boundary growth condition, implying an $O(N^{-1+\epsilon})$ RMSE for scrambled nets. We then establish these conditions for diffusion probability-flow transport under mild bounded-derivative assumptions on the learned vector field, explicitly controlling the boundary singularities introduced by the inverse Gaussian CDF. Experiments range from a 2D mixture to 784D images and a 40,960D conditional vorticity-assimilation task; in the latter, blocked scrambled Sobol'sampling reduces the randomization standard deviation of nonlinear accuracy metrics at essentially unchanged online denoising cost. Together, these results give a theoretical and empirical foundation for combining diffusion generative modeling with high-precision RQMC integration.

Jianlong Chen, Yifeng Yu · 0 citations
Preprint Aug 2026

Exploiting Exact Conditionals Improves Conditioning: Provably Fast Mixing Time Bounds By Sampling from the Marginal

The problem of sampling from a probability distribution arises in many applications such as posterior sampling in hierarchical Bayesian inverse problems and Gaussian processes for machine learning. Markov chain Monte Carlo (MCMC) algorithms are often used for sampling from a target probability distribution, but implementations can be computationally expensive, especially for large-scale problems. In certain applications, the target distribution naturally factorizes into a lower dimensional marginal distribution and a conditional distribution that allows exact sampling. We describe an MCMC algorithm called MarCo that exploits such a structure and generates a Markov chain via Metropolis-Hastings sampling from the marginal distribution, followed by sampling from the exact conditional distribution. By design, MarCo constructs a Markov chain on the joint space that inherits the convergence behavior of the marginal MCMC algorithm. This provides multiple theoretical and computational advantages. We prove that MarCo can achieve improved mixing time upper bounds compared to direct sampling from the joint distribution. Moreover, compared to one-block methods that also exploit marginal-conditional structure, we use the framework of Peskun-Tierney ordering to show that MarCo has a larger right spectral gap and smaller asymptotic variance, thus leading to superior convergence properties. Numerical results illustrate the performance benefits of MarCo and are provided for various problems, including a semi-blind image deblurring example.

Abhijit Chowdhary, Federica Milinanni, Julianne Chung et al. · 0 citations
Preprint Aug 2026

Blue Noise as a Lattice Gibbs Ensemble

Blue-noise sampling is widely used in computer graphics, but existing methods separate statistical modeling from scalable generation. Optimization and transport methods produce high-quality point sets by coupling all samples together. Procedural and tile-based samplers are local, but define their output only implicitly. We formulate blue-noise generation as sampling from a Gibbs distribution over binary lattice occupancies with pairwise repulsive interactions. Density, repulsion strength, interaction scale, and kernel hardness are parameters of this distribution. Because the energy sums over pairs, distant interactions can be dropped with a bounded change to the distribution, leaving a Markov random field of bounded degree. To sample it, we trace the Markov chain backward from the state we want, following Coupling Towards The Past, and cut the trace at a fixed depth. This bounds the cost, and it bounds the region each sample depends on. A tile generated on its own, with a sufficient halo, is then bit-identical to the same region generated on any larger domain, in any order and with no communication between tiles. Memory is set by the tile size, not by the output size, and accuracy is traded against cost through parameters with a proven error bound rather than by switching algorithms. We validate the model, the sampler, and these guarantees separately. The ensemble reproduces standard blue-noise spectra and moves continuously between them as its parameters vary. The sampler matches its predicted work and memory. Tiled output is verified bit-identical to full-domain generation. We demonstrate adaptive stippling at 14K, where existing methods need memory proportional to the output, along with multi-class extensions.

Zhu Yi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.