While stochastic diffusion samplers such as DDPM better preserve the enstrophy spectrum during rollouts in the stochastic setting, deterministic samplers such as DDIM and DPM-2 show better spectral preservation in the deterministic setting.
Abstract
We benchmark transport-based generative models as well as distillation-based few-step methods for the probabilistic forecasting of stochastic fluid flows, with a particular focus on performance under limited inference budgets. All methods are evaluated on a two-dimensional Kolmogorov flow with stochastic forcing. We measure one-step distributional accuracy against large simulated reference ensembles and assess whether the invariant measure is preserved during autoregressive rollouts via the enstrophy spectrum. On the stochastic task, flow matching achieves the most accurate one-step conditional distribution at high inference budgets, while the second-order exponential integrator DPM-2 is strongest at very low NFE. Few-step distillation methods are competitive with the multi-step methods and preserve the enstrophy spectrum particularly well. A deterministic control task, in which the forcing over the prediction interval is observed, separates aleatoric from epistemic uncertainty. Model performance does not translate between the two settings: the distilled models are competitive on the stochastic task but least accurate on the control task. While stochastic diffusion samplers such as DDPM better preserve the enstrophy spectrum during rollouts in the stochastic setting, deterministic samplers such as DDIM and DPM-2 show better spectral preservation in the deterministic setting.
This work exploits the affine state update to obtain the exact one-step conditional-mean sensitivity by differentiating normalized reaction propensities, and defines the propensity straight-through (PST) estimator, a temperature- and Gumbel-free path to scalable gradient-based learning through exact stochastic trajectories.
This work proposes Teacher Rollout Extension (TREX), a knowledge distillation framework that transfers the predictive capability of a pretrained foundation model into a compact and efficient student that can match or surpass the teacher's accuracy while reducing the number of parameters by several orders of magnitude and achieving more than an order-of-magnitude speedup in inference.
Daniel Musekamp, Boshra Ariguib, Andrei Manolache et al.· 0 citations
This work proposes Neural Kolmogorov Equations (NKEs), a deterministic, infinite-dimensional reformulation of Neural SDEs based on the Kolmogorov Forward equation, transforming the learning problem from modelling individual stochastic trajectories to modelling the evolution of probability densities.
This work composition the observables'likelihoods in per-task free-routed last-layer beliefs on a shared backbone absorbs unit-dependent loss scaling into likelihood parameters learned in the same gradient pass, and results land where theory puts them.
LatentFlow is introduced, a single framework for conditioning stochastic processes, with no learned neural approximations and no training, that enables conditional sampling in seconds on a single desktop CPU across model classes that have never shared a scalable method.
Louis Sharrock, L. Astfalck, Henry B. Moss· arXiv.org· 0 citations
A central question in physical inference is whether strongly constrained dynamical systems can realize accurate input--output maps through their own finite-time evolution. We study this question in Kuramoto phase networks, whose deterministic dynamics form an input-conditioned gradient flow and whose predictions are read directly from output oscillators. As a constructive training approach, we develop a two-stage teacher--student procedure. A neural teacher is first converted into an explicit phase trajectory whose terminal oscillator activations reproduce the teacher outputs, and the Kuramoto parameters are trained by matching the student vector field along this prescribed path. Because accurate teacher-forced path matching does not ensure accurate autonomous inference, we then differentiate through the autonomous finite-time rollout and directly align its terminal output with the neural target. The resulting oscillator system, with $74$ oscillators, reaches mean test accuracies of $96.711\%$ on MNIST and $86.399\%$ on Fashion-MNIST. This capability persists across neural-teacher architectures, matched system sizes, thermal perturbations, and integration-grid refinement. Together, these results provide a constructive demonstration that a strongly constrained, small-sized Kuramoto gradient-flow system can be trained for high-accuracy finite-time inference through a direct oscillator readout.
Yi Cheng, Zong-Li Lin· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.