Skip to content
Preprint

NestyNet. II. Coherent Function-Space Posteriors from Scientific Neural Surrogates (or How to Avoid Expensive MCMC)

Aug 2026 · 0 citations · 5 references
Physics

TL;DR

A low-dimensional posterior for the fitted function itself, for scientific neural surrogates trained with second-order optimization, provides a highly efficient route to uncertainty propagation for derivative-dependent scientific inference.

Abstract

Scientific analyses increasingly use flexible neural networks, but their thousands of correlated parameters make it challenging to interpret the associated uncertainties. Here we develop a low-dimensional posterior for the fitted function itself, for scientific neural surrogates trained with second-order optimization. Linearizing the fitting procedure with respect to the randomized residual rows gives a measurement-to-function transport, the linear map, assembled from the converged Jacobians and Gauss--Newton curvature, that carries measurement perturbations into the function perturbations that refitting would produce. Its leading singular functions define coherent deformation modes. Independent Gaussian coefficients then generate smooth function draws, so that any derived quantity, including those requiring derivatives or integrals of the draw, inherits the posterior. The construction distinguishes repeated-experiment covariance from the local Gauss--Newton/Laplace posterior and propagates both to correlated quantities of scientific interest. The result is conditional on the fit's declared choices (architecture, hyperparameters, active set, and optimization branch), and every fit is certified as converged by checking that a further optimization step would change the fitted predictions by less than a chosen small fraction of the measurement errors. Our primary example is an 800-parameter phase-space distribution function fit for a mock stellar disk. Four uncertainty coordinates, two orders of magnitude fewer than the fitted parameters and stable under refinement of the force basis, capture 99\% of the vertical-force posterior variance, and 4000 coherent draws propagate through the force, total-density, surface-density, and frequency calculations in 0.8s. The method provides a highly efficient route to uncertainty propagation for derivative-dependent scientific inference.

View source

Similar papers

Preprint Aug 2026

NestyNet. I. Physics Functions Are Hard to Fit with Neural Networks: A Framework for Accurate Surrogates and Analytic Derivatives

Many of the smooth functions that matter most in physics are precisely the ones that standard neural network methods struggle to fit accurately. Here we present NestyNet, a coupled model-and-optimizer framework capable of fitting such targets to high accuracy while also delivering their gradients, Hessians, Laplacians, and antiderivatives analytically and at low cost. This makes it a natural substrate for scientific machine learning tasks. The model is a deterministic segmented analytic surrogate, and its optimizer is a second-order Levenberg--Marquardt scheme whose damping and linear solves are tailored to the stiff, strongly correlated parameter geometries induced by multiscale and sharply structured targets typical in physics and other scientific applications. On the AI Feynman benchmark of 120 physics equations, NestyNet achieves median improvement factors of $2\,100\times$ for function values, $1\,400\times$ for first derivatives, and $780\times$ for second derivatives relative to standard neural networks trained with first-order optimization (Adam). Even after refining those fits with quasi-Newton (L-BFGS) optimization, the corresponding improvements are $540\times$, $450\times$, and $250\times$. Owing to the analytic design it is up to $\approx 44\times$ faster than vectorized automatic-differentiation (autograd) baselines, with a margin growing with model size. The same analytic-derivative framework also supports vector- and complex-valued targets, measurement uncertainties in both inputs and outputs, and constraints, and its modules can be composed flexibly to build scientifically useful model architectures, all without reverting to autograd. Together, these components provide a practical modular framework for fitting difficult scientific surrogates while delivering accurate differential operators for subsequent analysis.

Rodrigo Ibata, Wassim Tenachi, F. Diakogiannis et al. · 0 citations
Review Aug 2026

NestyNet. III. Symbolic Regression from Analytic Neural Surrogates

Many physical laws are simple only after the right representation, decomposition or internal coordinate has been found, but discovering that structure from data is combinatorially hard. This task is symbolic regression (SR), the search for closed-form expressions that fit data without assuming a fixed model class. Here we present NestyNet-SR. A neural surrogate with analytic derivatives is used to detect separability, recursively reducing multivariate problems to simpler neural atoms. These atoms are distilled into closed form by a tiered symbolic-search stack, whose final tier is a novel factorized symbolic search that separates structure from calibration. Composing candidate internal coordinates freely, it scores each coordinate by how well calibrated functions of it (e.g., polynomials, power laws, sinusoids) fit the data, so the constants of those calibrated maps, however deeply nested in the final expression, are fitted rather than searched. The method supports multi-dataset regression, automated feature discovery, and dimensional-analysis pruning. On the SRBench AI~Feynman benchmark, NestyNet-SR achieves exact symbolic recovery of all 120 noiseless equations, the first such result, and under noise a statistical audit certifies which structures survive. As a real-data vignette, given only the separate mass-model components of SPARC-survey galaxies, the algorithm discovers the baryonic acceleration coordinate, reproduces the established mass-to-light and acceleration scales and the non-unique form of the radial acceleration relation, and adds held-out-galaxy generalization, a calibrated symmetry abstention, and a posterior for the local slope of the law. Analytic derivatives thus provide a practical route from neural surrogates to interpretable closed-form empirical laws.

Rodrigo Ibata, Wassim Tenachi, F. Diakogiannis et al. · 0 citations
Preprint Aug 2026

NestyNet. IV. Laws Chosen by Nothing in Advance

Differential-equation (DE) discovery tends to break down precisely where much of physics begins. Fields are coupled, governing laws are nonlinear in the state, amplitudes, coordinates, or operators of interest, yet derivatives must remain consistent across fields, channels, and differentiation orders. NestyNet-DE addresses this via two complementary DE search strategies, both employing analytic derivatives from segmented neural surrogates: a sparse-library route linear in the outer coefficients, and a new operator-factorized route that searches directly over equation structure rather than a fixed library, recovering compositional laws the sparse route misses. The recovered laws can be strongly nonlinear in states, fields, coordinates, and their couplings. The framework also handles multi-dataset shared-support discovery, complex and vector laws, and Hamiltonian discovery from phase-space trajectories. Beyond a law's form, the same data yield its geometry, the Lie point symmetries of the recovered equation. We apply the pipeline to 30 years of daily ephemerides of $308$ main-belt asteroids and recover the reduced-Kepler hierarchy (areal law, inverse-square force, and reduced Hamiltonian), where a discovered rotational symmetry fixes the centrifugal coefficient rather than fitting it. We also present a benchmark of 57 real-valued ODEs and 26 complex-valued systems, including the Schr\"odinger and Dirac equations, with Maxwell's equations as a coupled vector-PDE case study. Here, discovery is not shackled to a fixed library, nor does it end at the equation. It composes laws that no dictionary anticipated, and closes the loop from data, to a law chosen by nothing in advance, to the geometry that explains and integrates it.

Rodrigo Ibata, Wassim Tenachi, F. Diakogiannis et al. · 0 citations
Jul 2026

ANGLE: Angular Neural Generative Learning via Engression

Circular data, representing angles or directions, are frequently encountered in computer vision, biology, geology, and meteorology. Traditional regression targets the conditional mean, which is often geometrically misleading for circular responses under multimodal, skewed, or asymmetric data structures. To address these limitations, a lightweight deep generative framework, namely ANGLE, is introduced for non-parametric distributional regression on the circle. The full conditional distribution of an angular response, given Euclidean and circular covariates, is learned through a generative map optimized via a generalized circular energy score (GCES) loss. Desirable theoretical properties, including the strict propriety of the loss and the rotational equivariance of the estimators, are established. Furthermore, both pre- and post-additive noise models are accommodated. A unified toolbox is provided for advancing previously underexplored challenges in circular statistics: extrapolation, sufficient dimension reduction, and conditional distribution equality testing. The framework's efficacy is demonstrated through extensive simulations and real-world applications. Specifically, the proposal is utilized for object pose estimation from imagery and wind direction prediction, which are integral to surveillance, autonomous vehicles, and energy systems, respectively. Superior predictive performance and robust uncertainty quantification of the proposed method in these tasks are revealed.

Rajdeep Pathak, Archi Roy, Tanujit Chakraborty · 0 citations
Preprint Aug 2026

Quantification and Decomposition of Uncertainty Using Sliced-Normal Distribution: With Applications to NASA Data

Modeling multivariate distributions with nonlinear dependence, multimodality, and tractable analytical structure for downstream applications is a central challenge in uncertainty quantification. Sliced Normal (SN) distributions were introduced in prior works at the National Aeronautics and Space Administration (NASA) to address this need by representing densities through polynomial feature maps. This construction provides a compact algebraic alternative to more opaque generative models, while retaining the ability to capture nonlinear parameter dependencies and multi-modal behavior. In this paper, we build on the SN framework and develop several improvements that make the approach more reliable and scalable. First, we reformulate SN parameter estimation as a convex optimization problem over a positive semidefinite matrix, replacing the original nonconvex likelihood search with a formulation amenable to standard optimization tools. Second, we clarify the expressive power of the SN class by connecting polynomial log-density modeling to a Stone--Weierstrass-type universal approximation argument on compact domains. Third, we propose a high-dimensional fitting procedure that partitions variables into approximately independent groups, fits SN models within each subgroup, and then assembles the subgroup models through a cross-block completion step to recover residual dependence. We demonstrate the resulting SN modeling pipeline on NASA loss-of-control flight data, where the method captures nonlinear dependence patterns in both low-dimensional slices and a higher-dimensional block-assembled model.

Arindam Roychowdhury, Luis G. Crespo, H. Lam · 0 citations
Preprint Aug 2026

On the Principles Behind Neural Network Optimizers

Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. This thesis develops a principled grounding for Adam and motivates new designs. First, we revisit Adam's divergence--convergence debate and show the existence of a problem-dependent phase transition: with properly chosen, batch-size-dependent hyperparameters, Adam converges, whereas under small-$\beta_2$ regimes it can diverge. Second, we investigate why Adam substantially outperforms SGD on Transformers through Hessian structure. We find that the Hessian evolves toward a near-block-diagonal form along training, accompanied by strong block heterogeneity. We prove that this structure makes Adam's diagonal preconditioner effective. We further show that this special Hessian structure originates from consecutive multiplications of large matrix variables, and we provide a rigorous analysis based on random matrix theory. Finally, these insights motivate Adam-mini, a new optimizer that reduces Adam's memory footprint by 50\% while preserving its performance. Our results also have broader implications beyond Adam: they reveal new local structures in matrix-based nonconvex problems, and also help understand and improve recent NN optimizers, such as Muon.

Yushun Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.