Skip to content
#small language model Open access

Beyond Loss Convergence: The Geometric Convergence Index and the Crystallization Horizon of Transformer Manifolds

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Materials Science

Abstract

Updated version. Fore details see changelog infra. Abstract of Beyond Loss Convergence: The Geometric Convergence Index and the Crystallization Horizon of Transformer Manifolds Determining when Large Language Model (LLM) pre-training is structurally complete remains one of the most critical open challenges in artificial intelligence. Contemporary foundation model pipelines rely on open-loop stopping criteria, such as fixed token budgets derived from Chinchilla scaling laws or cross-entropy loss flattening. However, on complex, high-entropy datasets, scalar loss flatlines relatively early upon reaching the Epistemic Floor (ℒ_epistemic → 0), where reducible structural error is exhausted and scalar loss becomes dominated by irreducible aleatoric data entropy (ℒ_aleatoric). Standard open-loop evaluation interprets this flatline as optimization stagnation or overtraining, frequently halting runs prematurely. In this work, we introduce the Geometric Convergence Index (GCI), a closed-loop, physical stopping framework that audits the internal hidden-state manifold during training. By tracking six coordinate-disentangled information-geometric criteria, inter-head fiber curvature (σ²_head), structural differentiation velocity (σ̇²_head), attention entropy variance, governor control effort (Ẇ_gov), Pawula Fokker–Planck stability (R_Pawula), and macroscopic order (R_t), the GCI establishes an exact, physics-grounded threshold for representational completion: the Crystallization Horizon (GCI = 1.0 for 10,000 consecutive steps). Telemetry reveals that beneath flat scalar loss plateaus (ℒ ≈ 6.98 - 7.83 nats), the model actively executes Hidden Descent, carving specialized head topologies. Furthermore, during a massive 50,000-step distribution shift under the pH-Stirred Mixture Protocol, the GCI correctly tracked a transient drop to 0.00 (Early Plasticity) followed by a rapid recovery to 0.33 (Active Descent), validating its role as a dynamic seismograph for representational maturity rather than a static thermometer. GCI provides foundation model labs with a dynamic, hardware-native stopping standard that prevents both premature termination and multi-million-dollar compute waste. Introduction: The LLM Stopping Problem Under modern scaling paradigms, Large Language Models (LLMs) are trained on massive multi-trillion token datasets (chinchilla, vaswani2017). Determining the precise completion point of a pre-training run currently relies on two open-loop heuristics: Static Token Budgets: Prescribing a fixed compute budget according to Chinchilla empirical scaling laws (chinchilla). This treats optimization as an arbitrary calendar schedule rather than a dynamic physical process. Open-Loop Loss Convergence: Monitoring the scalar cross-entropy loss curve and initiating learning rate warmdown when loss reductions plateau. Both heuristics suffer from a fundamental theoretical flaw: open-loop scalar loss is an inadequate sufficient statistic for internal representation learning (slt, massif). Consequently, foundation model labs face a multi-million dollar optimization dilemma: terminating training prematurely at the Epistemic Floor leaves the model's internal routing topology uncrystallized and severely degrades downstream reasoning capabilities, while continuing compute expenditure blindly past the point of geometric saturation wastes massive amounts of GPU cluster energy and capital without yielding proportional representational gains. To resolve the LLM stopping problem, this paper introduces the Geometric Convergence Index (GCI), a closed-loop stopping criterion that monitors latent information geometry directly during training, defining the exact point of geometric stasis: the Crystallization Horizon. Theoretical Background: Latent Kinematics & Epistemic Floor Riemannian Trajectory Transport & SDE Kinematics Representation propagation through depth ℓ ∈ {1, …, L} is modeled as continuous-time trajectory transport along a Riemannian manifold ℳ induced by attention metric tensors g^A_ij(𝐡) (amari2016information, nielsen2020riemannian). Trajectory increments follow the Langevin Stochastic Differential Equation (SDE): d𝐳_t = 𝛍_drift(𝐳_t) dt + σ_diff d𝐖_t where 𝛍_drift is first-order Kramers–Moyal drift velocity and σ_diff = √D⁽²⁾_eff is diffusion noise amplitude (pawula1967). Dynamics operate in a drift-dominated Ballistic Flow regime when directional Signal-to-Noise Ratio satisfies ρ_dir = ‖𝛍_drift‖ / σ_diff ≫ 1. In this regime, individual token trajectories travel along deterministic geodesic highways, unconfounded by random diffusion noise (massif). The Epistemic Floor and Hidden Descent Under Singular Learning Theory (SLT) (slt, watanabe2009algebraic), parameter space contains singular varieties of functionally equivalent configurations (parameter fibers) (grillo2026). Cross-entropy loss decomposes into two distinct components: ℒ_total = ℒ_epistemic + ℒ_aleatoric Here, ℒ_epistemic represents reducible structural error (model ignorance of the underlying data manifold), while ℒ_aleatoric represents irreducible entropy inherent in the data distribution (e.g., syntactic noise or high-entropy multi-lingual tapestries). When a model successfully resolves the structural rules of the manifold, ℒ_epistemic → 0. Total loss ℒ_total then flatlines, bounded strictly by aleatoric noise. Standard early-stopping algorithms misinterpret this physical boundary as a lack of gradient signal, triggering premature termination. When the model reaches the Epistemic Floor, macroscopic knowledge acquisition ceases, but internal manifold optimization continues. As demonstrated by the MASSIF control framework (massif), attention heads desynchronize into differentiated geometric routing operators without altering cross-entropy loss. To distinguish between active structural specialization and true geometric completion, internal state variables must be audited directly. Bridging Static Parameter-space Geometry with Dynamic, Continuous-time Hidden-state Trajectory Execution The tree papers establish together a unified, self-regulating foundation model stack across three operational levels: Foundational Trajectory Telemetry & Operational Mappings: We investigate candidate operational correspondences connecting static algebraic-geometric properties under Singular Learning Theory (SLT) - parameter fibers, fiber volumes, and the Real Log Canonical Threshold (λ) - to dynamic differential-geometric telemetry on attention-induced Riemannian manifolds. These mappings are treated as operational hypotheses rather than established mathematical equivalences. Specifically, we propose that parameter redundancies (padding functions) may correspond operationally to the residual component of a coordinate-disentangled ANOVA hidden-state decomposition; that the macroscopic order parameter R_t may serve as an empirical probe for structural phase transitions; and that the Alpha Potential Well U(α) acts as a curvature-motivated regularizing potential against scaling symmetries. Kramers–Moyal expansions and Pawula-based diagnostics provide empirical support, but not proof, for a low-order drift–diffusion description, with higher-order diffusion terms remaining small relative to second-order variance Dʌ(4)≪(Dʌ(2)ʌ2). Extended high-resolution auditing of MyceliaLM (1.62B parameters) reveals Persistent Individual Ballistic Flow (PIBF), wherein token representations move along deterministic, drift-dominated geodesics (ρ_dir>100) while cross-token directional alignment remains near zero (c_≈0.000) and head variance spikes (σ_headʌ2≈0.730–0.760), consistent with Spontaneous Symmetry Breaking into a specialized geometric routing swarm. Across 238 snapshot audits, the relative variability of the Optimization Response Function χ_R exceeds that of directional SDE SNR ρ_dir by a factor of >148×, establishing Kinematic-Thermodynamic Decoupling as an empirical phenomenon. Furthermore, near-zero correlations between structural differentiation velocity (dσ_headʌ2/dt) and scalar loss shifts (ΔL≈0) motivate the Hidden Descent Hypothesis: internal representational specialization can proceed independently of scalar loss at the Epistemic Floor. This is the subject treated by the present paper, Beyond Scalar Loss: An Empirical Observatory for Persistent Ballistic Flow and Hidden Descent in Transformers. Closed-Loop Pre-Training Completion & Stopping Standard: To solve open-loop pre-training stopping ambiguity when scalar loss plateaus at the Epistemic Floor (L_total = L_epistemic + L_aleatoric), we introduce the Geometric Convergence Index (GCI). The GCI audits six coordinate-disentangled information-geometric criteria: inter-head fiber curvature saturation (σ²_head), differentiation velocity decay (dσ²_head/dt), attention entropy variance stability, governor work rate minimization (Ẇ_gov), Pawula Fokker–Planck stability (R_Pawula), and macroscopic order stability (R_t). The GCI defines a physics-grounded stopping standard, the Crystallization Horizon (GCI = 1.0 for 10,000 consecutive steps), that terminates pre-training the exact moment representational maturity is reached, preventing premature halt or multi-million-dollar compute waste. This is the subject of the paper, Beyond Loss Convergence: The Geometric Convergence Index and the Crystallization Horizon of Transformer Manifolds. In-Vivo Runtime Governance & Silicon Alignment: Extending forward-pass geometry monitoring to inference, we introduce the Kinematic Friction Rail and the Cache-Fused Kinematic Rail (CFKR). By monitoring depth-dependent hidden-state variance differentials via the Friction Delta (Δ = V_late - V_early), CFKR detects Geometric Flatlines (Δ → 0.00) in vivo prior to token emission, resolving the Permissive Consensus Paradox. Rather than relying on software-level additive logit bias matrices (M_ghost) and token eviction, CFKR executes ke

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.