Given a strongly convex function $u$, equip $R^d$ with a Riemannian metric given by the Hessian $\nabla^2 u$. This is a so-called Hessian manifold. Given a probability density $\mu$ one may run a Langevin diffusion intrinsic to the manifold with stationary distribution $\mu$. Such (Hessian) manifold-valued Langevin diffusions are called Mirror Langevin diffusions (MLD) which have recently become popular. One of the questions we explore is whether, given $\mu$, one can choose $u$ to get an exponential convergence to equilibrium for the MLD, especially if $\mu$ is not strongly log-concave. Our results are based on Lyapunov function methods and give sufficient conditions for a Poincar\'e or a log-Sobolev inequality to hold for the MLD. These, in turn, imply exponential convergence. We also introduce a Markov chain approximation to the MLD given by a two step Gibbs sampler with stationary distribution $\mu$. This Markov chain is a variant of the Sinkhorn Markov chain introduced in arXiv:2307.16421 that is conjectured to converge to a time-inhomogeneous generalization of the MLD. Under suitable assumptions, we prove that the Markov chain has a guaranteed convergence rate in $\chi^2$ that is consistent with the diffusion time scale. Our proofs are based on ideas from entropic optimal transport and strong data processing inequalities.
We consider the Langevin diffusion $dX_t = - \beta \nabla V(X_t) dt + \sqrt{2} dB_t$ for a general nonnegative real-analytic potential $V$ and a large parameter $\beta$. In the large-$\beta$ limit the process is confined to the zero set of $V$, assuming that it starts there. We derive a candidate limiting evolution on the zero set. To do so, the zero set is partitioned into strata according to a measure of local codimension known as the local learning coefficient and its multiplicity. It is then shown that the Dirichlet form associated with $X$ converges in a certain sense to a hierarchy of Dirichlet forms corresponding to a stochastic evolution on the strata. This evolution is strongly biased toward higher-dimensional, or"more singular", strata. This result is motivated by a question from Watanabe's singular learning theory regarding the learning dynamics of overparameterized statistical models and the generalization puzzle in deep learning. The result suggests a mechanism for the observation that stochastic gradient methods tend to be biased toward singular solutions that generalize well.
We study the long-time behavior of the Wasserstein gradient flow of the squared Maximum Mean Discrepancy (MMD) between a probability measure $\rho$ and a target measure $\mu$, where the underlying kernel is given by a Coulomb potential. For $L^\infty$ target densities $\mu$, we establish the existence of global weak solutions starting from arbitrary Borel probability measures and prove that the density $\rho_t$ belongs to $L^\infty$ for any $t>0$. We also show that the H\"older norm can grow exponentially in time. On the flat torus ${\mathbb{T}}^\mathsf{d}$, we prove a global metric PL inequality for every finite-Coulomb-energy source and nearly uniform target. For general bounded, uniformly positive targets, we prove exponential decay of the squared MMD without requiring a lower bound on the initial data, using a defective PL inequality. We also prove that the usual PL inequality may fail when the target vanishes only at one point and that, when $\mathsf{d}\ge2$, no PL constant can hold uniformly over all targets satisfying a prescribed lower bound. On ${\mathbb{R}}^\mathsf{d}$, for $\mathsf{d}\ge2$, under radial symmetry, source-support inclusion, and target-positivity assumptions, we establish a PL inequality and exponential convergence. On the unrestricted whole-space class, neither a multiplicative squared-MMD decay modulus uniform over the initial datum nor a global PL inequality can hold. Finally, in every dimension and in both spatial settings, we prove that every Lagrangian critical point coincides with the target when $(\rho-\mu)^+$ is absolutely continuous. In dimension two, the energy supplies uniform tightness. This implies that if our constructed solutions have finite energy at some positive time, then they converge to the target narrowly and strongly in negative-order Sobolev spaces.
Antonin Chodron de Courcel, Matthew Rosenzweig· 0 citations
We introduce a new class of uniformly ergodic MCMC algorithms, termed Diffeomorphic Contraction Sampler (DCS), and provide fast non-asymptotic mixing guarantees for DCS targeting distributions on $\R^d$ with arbitrarily heavy polynomial tails. DCS provides a solution to a well-known problem for MCMC samplers, which typically struggle with the combination of unbounded high-dimensional state space and vanishing gradients. The DCS pulls back a target on $\R^d$ onto a Euclidean ball $B(R)\subset\R^d$ and then samples from the transformed density on the convex set $B(R)$ via algorithms such as the Ball Walk, Hit-and-Run and others. A radial diffeomorphic contraction is chosen so that the pull-back density on $B(R)$ is bounded, implying uniform ergodicity for \textit{all} targets with a finite polynomial moment. Non-asymptotic bounds for DCS require stronger assumptions such as log-concavity of the pull-back density. In practice, this is achieved approximately by a preconditioned automorphism of the ball $B(R)$, tuned via Variational Inference. Numerical simulation tests demonstrate that the DCS outperforms significantly the No-U-Turns sampler on multi-dimensional heavy-tailed targets arising as real-world posteriors in PosteriorDB benchmark. DCS also numerically outperforms in high-dimensional examples recently developed spherical projection samplers for heavy-tailed target distributions.
We study homogeneous diffusion martingales evolving in a bounded state space $D=[a,b]$, where $a$ and $b$ are zeros of the diffusion coefficient. We call a process of the form $Z_t=\mathbb{E}[B\mid\mathcal{F}_t]$, with $B$ a Bernoulli random variable, a Bernoulli-Doob martingale. Our main results establish a complete equivalence: every such diffusion martingale is a Bernoulli-Doob martingale (Theorem 2) and, conversely, every continuous time-homogeneous Markov Bernoulli-Doob martingale on a Brownian filtration arises from such a diffusion (Theorem 3). The intuitive reason is that a bounded martingale has constant expectation while accumulating variance, so it converges to the maximum-variance distribution with given mean and range, namely the Bernoulli. We further show that this Bernoulli limit is truly asymptotic: for any fixed finite horizon $T$, the probability of not yet having reached the boundary is strictly positive (Theorem 4), even when the individual boundaries are accessible. We clarify the relationship between Feller's boundary classification, the pathwise SDE framework, and the martingale constraint, showing that the martingale property forces absorption at any attainable boundary. The theory is illustrated with the $\Phi$-martingale, the Jacobi martingale, and applications to credit-risk modelling.
We study the squared singular value spectrum of a non-square product of independent real Gaussian matrices, equivalently the feature covariance spectrum of a deep linear neural network at initialization. Starting from the fixed-$m$ covariance diffusion previously obtained in the proportional depth-width limit, we record an equivalent matrix realization, describe its affine invariance, and derive the interacting diffusion satisfied by its eigenvalues. We then take a second limit, sending $m\to\infty$ on the accelerated spectral clock $\tau=mt$, which corresponds in this sequential construction to the relation $dm/n\to\bar\tau$. We establish convergence of the empirical spectral measure path to a deterministic mean-field limit and derive a closed Burgers equation for its $T$-transform. Together with the proportional depth-width limit, these results give a rigorous sequential route from the deep non-square Gaussian product to the free log-normal limit of its feature covariance spectrum; for more general initial laws, the transform yields a free multiplicative convolution form. We further analyze the support of the free log-normal law, give a fixed point iteration for numerical evaluation and a formal Marchenko--Pastur approximation at small time, and use the limiting spectrum to predict the risk in a toy random feature model.
Mufan Bill Li, Jaume de Dios Pont, M. Nica et al.· 0 citations
We establish the sharp logarithmic order $(\log N)^{-1/2}$ for the expected $p$-Wasserstein distance, induced by the supremum norm, between the empirical law of $N$ independent copies of a continuous It\^o process and their common path law. We only assume that the initial condition and the drift and diffusion integrands are controlled by a time-uniform random upper bound with a finite $\rho$-moment for some $\rho>p\geq1$. Under this assumption, we use an adaptive random time interval partition argument, which leads to a $(\log n)^{-1/2}$ functional quantization rate. A general transfer principle then converts the quantization estimate into a mean estimate and nonasymptotic deviation bounds for equal-weight empirical laws. Applications include empirical path-law estimates for path-dependent SDEs and a path-space propagation-of-chaos estimate for path-dependent McKean--Vlasov interacting particle systems.