Skip to content
#diffusion models Open access

FROM FIBERS TO FLOWS: RESOLVING THE THERMODYNAMICS OF NEURAL OPTIMIZATION IN RELU NETWORKS

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)
Plant and Biological Electrophysiology Studies

Abstract

ABSTRACT The ILIAD research programme poses a central question: do ReLU networks possess an intrinsic simplicity bias analogous to Solomonoff induction, and can this bias be characterized through the geometry of the fiber? We argue that this question admits a complementary formulation in the language of continuous-time stochastic differential equations (SDEs) on attention-induced Riemannian manifolds. We present three formal structural mappings between ILIAD's static algebraic-geometric framework (toric encoding maps, fiber volumes, and the Real Log Canonical Threshold) and our continuous-time trajectory telemetry framework (the coordinate-disentangled ANOVA hidden state decomposition, the macroscopic order parameter R_t, and the Alpha Potential Well metric scaling regulator). Rather than claiming mathematical identity, we model the Real Log Canonical Threshold (RLCT, λ) and the Constructive-Compensatory Ratio (R_t) as a conjectured phenomenological correspondence, where R_t acts as a live thermodynamic thermometer of representational complexity. We provide empirical evidence from the closed-loop training of Mycelia-LM, a de novo 1.62-billion-parameter self-governing transformer. We document a critical singular phase transition (grokking) where the Optimization Response Function χ_R = ∂ℒ_loss / ∂R_t flips sign. Furthermore, we expose a failure mode mapping Goodhart’s Law directly onto the geometric control plane: a proxy-gaming loop where the Lesson-Based Retrieval (LBR) engine over-inflated attention target norms beyond physical manifold capacity, triggering an AlphaScale safety release and catastrophic coordinate collapse (R_t: 0.638 → 0.122 in 250 steps). We resolve this ex-ante by implementing context-level soft target correction (ctx[\text{"alpha_norm_target"}] = \text{realistic}), proving that: continuous-time geometric telemetry can predict, prevent, and stabilize representation failures in self-governing systems. 1. INTRODUCTION: TWO LANGUAGES FOR ONE PHENOMENON The ILIAD programme identifies a fundamental gap in AI safety: the absence of a sufficiently general, predictive mathematical account of neural network generalization. Their proposed resolution proceeds through algebraic geometry, specifically, through the toric encoding map Φ_θ, which maps d-dimensional real parameter space to d'-dimensional real space and reveals the tractable geometric structure of ReLU networks. This structure includes the fiber F of f, defined as the set of all parameter configurations θ for which the toric encoding map produces identical functional behaviour f. Our research programme, the Multiscale Attractor Stability and Stress Inference Framework (MASSIF), proceeds from a complementary direction. Rather than characterizing the static algebraic structure of the parameter space, we measure the dynamic trajectory of optimization as it traverses that space during the forward pass. The MASSIF framework equips a transformer with real-time differential-geometric telemetry: a 13-observable kinematic state vector, tracking velocity, curvature, torsion, and Lyapunov divergence of the hidden-state trajectory 𝐡𝐭 as it evolves on an attention-induced Riemannian manifold 𝐌 equipped with a covariant metric tensor 𝘨ᴬᵢⱼ. These two programmes are not competing. They are dual descriptions of the same underlying physical phenomenon: ILIAD asks: What is the geometry of the space through which learning moves? MASSIF asks: What is the geometry of the movement through that space? The former is a question about the manifold; the latter is a question about the geodesic. Both require differential geometry. Both produce safety-relevant, active observables. And both converge on the same empirical phenomenon: the phase transition from compensatory, high-dimensional brute-force memorization to constructive, low-dimensional structural generalization, what the broader literature calls grokking. 2. THE BRIDGE: THREE STRUCTURAL MAPPINGS To formally unite these dual descriptions, we establish three explicit correspondences connecting static algebraic topography with active, continuous-time trajectory telemetry: Mapping 1: Fiber Degeneracies ↔ ANOVA Residuals ILIAD Algebraic Concept: The parameter fiber F(f) representing all parameter configurations θ yielding identical functional output f. Larger fiber volumes represent higher parameter redundancy. MASSIF Dynamic observable: The squared Euclidean norm of the ANOVA residual vector: ‖resid_{c,t}^{(ℓ)}‖₂². The hidden-state residual stream h𝄂,t⁽ℓ⁾ at layer ℓ, sequence position t, and context c, is projected onto coordinate-disentangled subspaces: h_{c,t}^{(ℓ)} = μ^{(ℓ)} + 𝐩𝐨𝐬 t^{(ℓ)} + 𝐜𝐭𝐱 c^{(ℓ)} + 𝐫𝐞𝐬𝐢𝐝 {c,t}^{(ℓ)} By tracking the mutual incoherence between the positional basis and the context basis, we establish a dynamic, online proxy for near fiber volume: incoherence max_{t, c} ⟨ 𝐩𝐨𝐬_t / 𝐩𝐨𝐬_t , 𝐜𝐭𝐱_c / 𝐜𝐭𝐱_c ⟩ When incoherence drops below 0.15, the coordinate subspaces are cleanly orthogonalized, indicating a minimal parameter representation near an irreducible branch of the fiber. Mapping 2: Real Log Canonical Threshold ↔ Order Parameter (R_t) ILIAD Algebraic Concept: The Real Log Canonical Threshold (RLCT, λ). In Singular Learning Theory, λ is the fundamental complexity measure governing the asymptotic rate at which the Bayesian posterior concentrates. MASSIF Dynamic Observable: The macroscopic order parameter, the Constructive-Compensatory Ratio (R_t), which acts as a live, empirical probe for RLCT transitions: R_t = Π_α(t) / (Π_FFN(t) + Π_MPC(t)) where Π_α, Π_FFN, and Π_MPC represent the optimization pressures (stress tensors) exerted by the attention, feed-forward, and model predictive control governors, respectively. Phenomenological Correspondence: Rather than claiming mathematical identity, we model the RLCT (λ) and the order parameter (R_t) as a conjectured monotonic correspondence across phase boundaries. In the Compensatory Regime (R_t < 0.8), the system is trapped in a highly degenerate, high-volume variety carrying extensive padding (memorization). At the critical point (R_t ≈ 1.2), the Optimization Response Function 𝜒_𝑅 = ∂ℒ_loss / ∂𝑅_𝑡 decouples (𝜒_𝑅 ≈ 0). In the Constructive Regime (R_t > 1.5), 𝜒_𝑅 flips negative, and the model condenses onto a lower-dimensional, highly generalized manifold. Mapping 3: Toric Padding ↔ The Alpha Potential Well ILIAD Algebraic Concept: Toric Padding, the algebraic symmetries and scaling redundancies (e.g., group actions) that artificially inflate fiber volume without changing model behavior. MASSIF Dynamic Observable: The Alpha Potential Well, a symmetric quadratic loss term defined over layers 1 … L: U(α) = λwell ∑ℓ=1L [ (αattn(ℓ) - 1.0)² + (αffn(ℓ) - 1.0)² ] Mathematical Grounding: Under the weight-decay-free Manual Muon optimizer, U(α) acts as an effective metric scaling constraint. Because self-attention weights directly define the localized attention-induced Riemannian metric tensor 𝘨ᴬᵢⱼ, penalizing deviations of αattn from unity conformally scales the metric. This prevents unchecked scaling degeneracies from blowing up or collapsing local geodesic ball volumes, acting as an effective curvature modulator. 3. CONTINUOUS-TIME TRAJECTORY KINEMATICS & SDE VALIDATION To model the autoregressive uncertainty of inference, the latent trajectory 𝐳_𝐭 is formulated as an Itô Stochastic Differential Equation (SDE) driven by a semantic velocity field and Brownian noise: dzₜ = v_logic(zₜ)dt + σ dWₜ To rigorously validate whether the empirical hidden states of Mycelia-LM exhibit the statistical properties of a true SDE with approximately separable drift and diffusion, we propose the Kramers-Moyal Empirical Verification Protocol. By extracting the transition probability P(zₜ₊Δₜ | zₜ)σ dWₜ across adjacent layer steps (Δt = Δℓ = 1)ₜ we compute the empirical drift vector 𝐃⁽¹⁾ and diffusion tensor 𝐃⁽²⁾: 𝐃⁽¹⁾(𝐳) = lim_{Δt → 0} 1/Δt 𝔼[𝐳_{t+Δt} − 𝐳_t | 𝐳_t = 𝐳]𝐃⁽²⁾(𝐳) = lim_{Δt → 0} 1/(2Δt) 𝔼[(𝐳_{t+Δt} − 𝐳_t)(𝐳_{t+Δt} − 𝐳_t)ᵀ | 𝐳_t = 𝐳] By evaluating Pawula's Theorem, verifying that the fourth-order Kramers-Moyal coefficient is negligible 𝐃⁽⁴⁾ → 0 we can mathematically prove if the latent transitions behave as a continuous drift-diffusion process or a jump process. Under SDE validation, dynamics cleanly bifurcate into two regimes: Regime I (Coherent Reasoning): Under logical dominance (SNR \rho = |\mathbf{v}_{\text{logic}}|/\sigma \gg 1), expected net displacement |\mathbf{z}_T - \mathbf{z}_0|_2 scales linearly with sequence length T [82]. Curvature approaches zero, defining ballistic, directed geodesics. Regime II (Stochastic Collapse): Under noise dominance (ρ ≪ 1), net displacement scales sub-linearly as O(√T). Expected curvature stabilizes near unity, trapping intermediate representations in high-curvature, repetitive "Hesitation Loops". 4. ACTIVE EX-ANTE GOVERNANCE & GHOST PARAMETER RECOVERY Rather than diagnosing representing collapses post-hoc, MASSIF utilizes Fibonacci Coherence Attenuation to damp out high-frequency coordinate noise in the residual stream: α_atten⁽ℓ⁾ = γ⁽ℓ⁾ · exp(-I⁽ℓ⁾ · ℓ/L). The baseline attenuation multiplier γ⁽ℓ⁾ is modeled using successive ratios of the Fibonacci sequence, asymptotically approaching the Golden Ratio (ϕ⁻¹ ≈ 0.618033), ensuring that deeper layers undergo progressively heavier structural damping when local incoherence I⁽ℓ⁾ spikes. Furthermore, our closed-loop logs on Mycelia-LM (1.62B parameters) exposed a classic Goodhart's Law failure mode. The LBR engine, observing that raising the attention scaling target α_norm_target historically improved R_t, aggressively over-inflated the target to 60. Bounded by the localized Riemannian metric, the actual attention heads physically could not exceed a norm of ~30. Detecting this massive deficit, the AlphaScale governor relaxed the potential well to 0.979, killing the manifold compression. This caused the at

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Conference Sep 2010

Exploring the Sources of Waste in Kanban Software Development Projects

The application of agile software methods and more recently the integration of Lean practices contribute to the trend of continuous improvement in the software industry. One such area warranting proper empirical evidence is a project’s operational efficiency when using the Kanban method. This short paper takes a new angle and explores waste in the Kanban-driven software development project context. A preliminary research model is presented for helping the consequent replication of the study. The results from the empirical analysis suggest Kanban can be an effective method in visualizing and organizing the current work, but does not prevent waste from creeping in, although the overall project outcome may be successful.

Marko Ikonen, Petri Kettunen, Nilay V. Oza et al. · 67 citations · ⚡9

Related blog posts

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

GPT-Lab Sep 10, 2026

Responsible AI Must Consider Its Afterlife

AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.