A Statistical-Signature Framework for Detecting Irreducible Stochasticity in CFD Simulations
Abstract
Summary This preprint presents a 2D proof of concept for deciding, before an expensive simulation campaign, whether a pointwise (trajectory-level) grid-convergence study is an appropriate verification target. In chaotic flows, small differences, including those introduced by discretization, grow until pointwise agreement between grids is lost. Grid convergence can then hold for statistics but not for trajectories. The workflow has two modules: Module A (early warning). A coarse-grid twin run, in which one copy is perturbed by a relative amplitude of 10⁻⁶, measures the finite-time growth rate of the perturbation. A run is flagged if the perturbation reaches 10⁻² within 100 time units at a fitted growth rate above 0.1. A flag means pointwise agreement between grids can hold only up to a finite horizon. A proposed third outcome, "undecided", covers runs that grow but do not reach the threshold within the observation window. Module B (two-grid statistical agreement). Independent ensembles on two grids are sampled sequentially until the 95% confidence interval of the relative difference in a mean quantity is narrower than ±2%. The result is reported as a bound δ*. Main results (2D Kolmogorov flow, jax-cfd, pre-registered decision rules) The 64² twin test reproduced the 128² chaotic/non-chaotic classification in all 31 regime–initial-condition cases tested. In an 11-point viscosity sweep, both grids placed the onset of chaos at ν_c ≈ 0.031. The 64² sweep cost about a quarter of the 128² sweep. Module B bounded the 64²–128² difference in mean enstrophy at δ* = 4.9% (ν = 0.001) and δ* = 2.7% (ν = 0.005). In these two regimes, the ensemble size predicted after the first round was within 15% of the realized size. An analysis shows that the warning time is set by ln(δ_th/ε)/λ. This explains why the warning is slow near onset and gives a design rule for the observation window. Limitations Only the binary flag agreed across grids. The magnitude of the growth rate did not: a pre-registered tolerance failed in 7 of 14 cases. This is consistent with scale-dependent error growth reported in the literature. Both grids resolved the flow in the onset test. Whether an under-resolved coarse grid can anticipate fine-grid behavior is untested. The onset sweep used two initial conditions per viscosity and a single perturbation amplitude. The work does not establish a general CFD grid-convergence criterion, and it does not show convergence to a continuum solution. No 3D or wall-bounded flows were tested. Development history Version 4.0 replaces the statistical-signature detector of Versions 1–3. That detector was falsified on public 3D turbulence data, where it rated genuine turbulence as safer than its own deterministic null (AUC = 0.11). The full record of failures is included. Reproducibility The accompanying package contains, for all seven test stages: pre-registration protocols scripts and result files a seed manifest, environment and version lists, run times, and SHA-256 checksums The onset-sweep results were regenerated from the registered code and seeds under a newer JAX release, and every reported value was reproduced. Follow-up tests (F1–F7) are stated as falsifiable predictions; the priorities are perturbation-amplitude robustness (F1) and an under-resolved coarse grid (F2). Disclosure Design, execution and evaluation were carried out by the author with an AI assistant (Claude, Anthropic). The code has not yet been reviewed or replicated by an independent CFD specialist.