Skip to content

1-Lipschitz Neural Networks on Hadamard Manifolds

Jul 2026 · arXiv.org · Vol abs/2607.19335 · 0 citations · 63 references
Mathematics Computer Science

TL;DR

This work constructs and analyzes a class of 1-Lipschitz neural networks on Hadamard manifolds, and shows improved results from the nonexpansive denoiser over static, data-only, and Log-Euclidean denoising baselines, and empirically test its convergence properties.

Abstract

Controlling the Lipschitz constant of a neural network is a standard way to promote robustness and stability. Most existing constraining strategies are designed for Euclidean spaces. In this work, we construct and analyze a class of 1-Lipschitz neural networks on Hadamard manifolds. Our layers are of gradient-descent type, $1$-Lipschitz, and quasi-$\alpha$-firmly nonexpansive. The core building blocks of the proposed architecture are Busemann functions, and we exploit the properties of Busemann gradient flows to design $1$-Lipschitz geometry-preserving layers. We provide explicit constructions and examples for hyperbolic manifolds and the manifold of symmetric positive definite (SPD) matrices. We test the proposed architecture in two numerical experiments: robust classification on the Poincar\'e disk and masked-Wishart covariance reconstruction. On the Poincar\'e disk, the proposed networks yield robust classifiers under hyperbolic perturbations. On the SPD manifold, we train SPD-valued denoisers and adopt them as a Plug-and-Play prior for a masked-Wishart covariance reconstruction problem. We show improved results from the nonexpansive denoiser over static, data-only, and Log-Euclidean denoising baselines, and empirically test its convergence properties.

View source

Similar papers

Sep 2026

The Geometric Whitney Problem and Approximations by Neural Networks on Manifolds

Abstract Why do neural networks overcome the curse of dimensionality? A common justification is that real-life high-dimensional data typically lie close to low-dimensional manifolds, and that neural networks can exploit this structure efficiently – overcoming the curse of dimensionality for their parameter counts. However, existing bounds depend on properties of the manifolds that cannot be read off from data alone. We close this gap. If a dataset locally looks like a low-dimensional linear space – a condition testable directly from data and derivable from the empirically supported manifold hypothesis under well-behaved conditions – then an approximating manifold M can be constructed. Neural networks can then approximate C1 functions uniformly on M, with parameter counts bounded purely in terms of computable properties of the data and overcoming the curse of dimensionality.

Jakob Konstantin Hecker · 0 citations
Preprint Sep 2026

Projected Gradient Method on Hadamard Manifolds

We study constrained smooth optimization problems on Hadamard manifolds with closed geodesically convex feasible sets. We analyze two projected gradient schemes: one with a constant stepsize and another with a backtracking line search. The constant-stepsize scheme is analyzed under the assumption that the objective function has a Lipschitz continuous Riemannian gradient, whereas the backtracking variant does not require this assumption to establish stationarity of accumulation points. For both schemes, we prove that every accumulation point of the generated sequence is first-order stationary under the respective assumptions, without requiring compactness of the feasible set; compactness is needed only to ensure the existence of accumulation points. When the objective function has a Lipschitz continuous Riemannian gradient, we derive iteration-complexity bounds of order \(O(1/\sqrt{N})\) for projection-based stationarity measures for both schemes, together with the corresponding \(\varepsilon\)-complexity estimates. For the backtracking scheme, the complexity analysis additionally requires the trial line-search stepsizes to be uniformly bounded away from zero. Under the same respective assumptions, the generated sequences are also asymptotically regular. Finally, we illustrate the practical performance of the methods by solving constrained Karcher mean problems on the manifold of symmetric positive definite matrices.

O. P. Ferreira, M. L. N. Gonçalves, Á. M. González et al. · 0 citations
Jul 2026

On the robustness of noisy solutions in non-convex neural networks

Using a finite energy message-passing algorithm, it is demonstrated numerically that thermal noise enables effective generalization in the regime of constraint densities where both recovering the teacher and finding a zero temperature solution are computationally hard.

Enrico M. Malatesta, A. Passalacqua, Riccardo Zecchina · 0 citations
Preprint Aug 2026

Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces

Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode an input function through point values on a fixed discretization. Building on the Topological DeepONet framework of Ismailov (arXiv:2603.11972), we replace point samples by continuous linear functionals drawn from the continuous dual of a Hausdorff locally convex space $({V},\{p_\alpha\}_{\alpha\in A})$, whose topology is generated by a point-separating family of seminorms rather than a single norm, and develop fixed and adaptive functional measurement systems. Measurements are combined with the coefficient-space Two-Step procedure of Lee and Shin (arXiv:2309.01020), while a training-only decoder and regularization stabilize the adaptive coordinates. We derive a discrete error decomposition separating measurement, output-basis, and neural-approximation errors, together with a Barron-rate refinement. The framework is evaluated on the antiderivative operator, a non-normable locally convex input space, heterogeneous Darcy flow, a controlled operator, and fixed-time and time-evolving Navier-Stokes vorticity operators. In the heterogeneous Darcy problem, the functional models retain nearly resolution-independent errors of 5.5-5.6% on unseen grids, while in the controlled problem adaptive measurements reduce the mean error below 1.2%. For the fixed-time Navier-Stokes problem, the Adaptive Topological DeepONet is the most accurate DeepONet-based model, attaining a mean relative $L^2$ error of 1.685% +/- 0.017% using 128 functional coordinates. A comparably sized Fourier neural operator (FNO; arXiv:2010.08895) achieves the lower error 0.832% +/- 0.172%, but requires the full 64x64 input field, twice the training time, and 10.7x greater peak GPU memory. The formulation provides compact, interpretable, and discretization-portable coordinates in the continuous dual $V'$, including for non-normable input spaces.

Khemraj Shukla, G. Karniadakis · 0 citations
Preprint Aug 2026

Is Grokking a Loss of Normal Hyperbolicity of the Interpolation Manifold?

A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm. In the language of dynamical systems, this is a fast-slow system in which the interpolation manifold plays the role of a slow manifold. We ask a question that this framing makes natural but the existing literature does not address: is the sharp generalization transition a loss of normal hyperbolicity of that manifold: a fold- or bifurcation-like event in which a normal restoring direction goes flat? Or does the manifold stay uniformly attracting while generalization happens by smooth drift? We propose a simple, optimizer-agnostic diagnostic: the smallest nonzero singular value $\sigma_{\min}^{+}(\mathbf J)$ of the residual Jacobian, which, for the squared loss, equals the slowest normal restoring rate of the manifold. On a two-layer ReLU network trained to grok modular addition under squared loss, $\sigma_{\min}^{+}(\mathbf J)$ does not collapse at the transition; it is near zero only before memorization and attains its largest values during the transition. The result holds across five seeds, and the six smallest singular values behave identically; there is no subspace-local collapse either. This is preliminary evidence against the bifurcation hypothesis and in favor of the smooth-contraction picture. We are explicit that a single-setting, gradual-transition experiment under Adam optimizer does not prove the absence of a bifurcation; it constrains where one could hide.

Suvinava Basak · 0 citations

Riemannian Deep Learning: Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of underlying geometries. It generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups and gyrogroups, and extends multinomial logistic regression from Euclidean space to SPD manifolds and then to general Riemannian manifolds. It further develops neural networks for several important geometric representations, including an unconstrained model of hyperbolic space, Busemann-based hyperbolic learning, and full-rank correlation matrices. Finally, it introduces adaptive and computationally efficient Riemannian metrics on SPD manifolds, including learnable Log-Euclidean geometries and fast, stable Cholesky-based geometries. The proposed methods are supported by theoretical analysis and validated through numerical experiments and applications in vision, signal processing, graph learning, and genomics.

Ziheng Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.