Beyond Loss Convergence: The Geometric Convergence Index and the Crystallization Horizon of Transformer Manifolds
Abstract
Updated version. Fore details see changelog infra. Abstract of Beyond Loss Convergence: The Geometric Convergence Index and the Crystallization Horizon of Transformer Manifolds Determining when Large Language Model (LLM) pre-training is structurally complete remains one of the most critical open challenges in artificial intelligence. Contemporary foundation model pipelines rely on open-loop stopping criteria, such as fixed token budgets derived from Chinchilla scaling laws or cross-entropy loss flattening. However, on complex, high-entropy datasets, scalar loss flatlines relatively early upon reaching the Epistemic Floor (ℒ_epistemic → 0), where reducible structural error is exhausted and scalar loss becomes dominated by irreducible aleatoric data entropy (ℒ_aleatoric). Standard open-loop evaluation interprets this flatline as optimization stagnation or overtraining, frequently halting runs prematurely. In this work, we introduce the Geometric Convergence Index (GCI), a closed-loop, physical stopping framework that audits the internal hidden-state manifold during training. By tracking six coordinate-disentangled information-geometric criteria, inter-head fiber curvature (σ²_head), structural differentiation velocity (σ̇²_head), attention entropy variance, governor control effort (Ẇ_gov), Pawula Fokker–Planck stability (R_Pawula), and macroscopic order (R_t), the GCI establishes an exact, physics-grounded threshold for representational completion: the Crystallization Horizon (GCI = 1.0 for 10,000 consecutive steps). Telemetry reveals that beneath flat scalar loss plateaus (ℒ ≈ 6.98 - 7.83 nats), the model actively executes Hidden Descent, carving specialized head topologies. Furthermore, during a massive 50,000-step distribution shift under the pH-Stirred Mixture Protocol, the GCI correctly tracked a transient drop to 0.00 (Early Plasticity) followed by a rapid recovery to 0.33 (Active Descent), validating its role as a dynamic seismograph for representational maturity rather than a static thermometer. GCI provides foundation model labs with a dynamic, hardware-native stopping standard that prevents both premature termination and multi-million-dollar compute waste. Introduction: The LLM Stopping Problem Under modern scaling paradigms, Large Language Models (LLMs) are trained on massive multi-trillion token datasets (chinchilla, vaswani2017). Determining the precise completion point of a pre-training run currently relies on two open-loop heuristics: Static Token Budgets: Prescribing a fixed compute budget according to Chinchilla empirical scaling laws (chinchilla). This treats optimization as an arbitrary calendar schedule rather than a dynamic physical process. Open-Loop Loss Convergence: Monitoring the scalar cross-entropy loss curve and initiating learning rate warmdown when loss reductions plateau. Both heuristics suffer from a fundamental theoretical flaw: open-loop scalar loss is an inadequate sufficient statistic for internal representation learning (slt, massif). Consequently, foundation model labs face a multi-million dollar optimization dilemma: terminating training prematurely at the Epistemic Floor leaves the model's internal routing topology uncrystallized and severely degrades downstream reasoning capabilities, while continuing compute expenditure blindly past the point of geometric saturation wastes massive amounts of GPU cluster energy and capital without yielding proportional representational gains. To resolve the LLM stopping problem, this paper introduces the Geometric Convergence Index (GCI), a closed-loop stopping criterion that monitors latent information geometry directly during training, defining the exact point of geometric stasis: the Crystallization Horizon. Theoretical Background: Latent Kinematics & Epistemic Floor Riemannian Trajectory Transport & SDE Kinematics Representation propagation through depth ℓ ∈ {1, …, L} is modeled as continuous-time trajectory transport along a Riemannian manifold ℳ induced by attention metric tensors g^A_ij(𝐡) (amari2016information, nielsen2020riemannian). Trajectory increments follow the Langevin Stochastic Differential Equation (SDE): d𝐳_t = 𝛍_drift(𝐳_t) dt + σ_diff d𝐖_t where 𝛍_drift is first-order Kramers–Moyal drift velocity and σ_diff = √D⁽²⁾_eff is diffusion noise amplitude (pawula1967). Dynamics operate in a drift-dominated Ballistic Flow regime when directional Signal-to-Noise Ratio satisfies ρ_dir = ‖𝛍_drift‖ / σ_diff ≫ 1. In this regime, individual token trajectories travel along deterministic geodesic highways, unconfounded by random diffusion noise (massif). The Epistemic Floor and Hidden Descent Under Singular Learning Theory (SLT) (slt, watanabe2009algebraic), parameter space contains singular varieties of functionally equivalent configurations (parameter fibers) (grillo2026). Cross-entropy loss decomposes into two distinct components: ℒ_total = ℒ_epistemic + ℒ_aleatoric Here, ℒ_epistemic represents reducible structural error (model ignorance of the underlying data manifold), while ℒ_aleatoric represents irreducible entropy inherent in the data distribution (e.g., syntactic noise or high-entropy multi-lingual tapestries). When a model successfully resolves the structural rules of the manifold, ℒ_epistemic → 0. Total loss ℒ_total then flatlines, bounded strictly by aleatoric noise. Standard early-stopping algorithms misinterpret this physical boundary as a lack of gradient signal, triggering premature termination. When the model reaches the Epistemic Floor, macroscopic knowledge acquisition ceases, but internal manifold optimization continues. As demonstrated by the MASSIF control framework (massif), attention heads desynchronize into differentiated geometric routing operators without altering cross-entropy loss. To distinguish between active structural specialization and true geometric completion, internal state variables must be audited directly. Bridging Static Parameter-space Geometry with Dynamic, Continuous-time Hidden-state Trajectory Execution The tree papers establish together a unified, self-regulating foundation model stack across three operational levels: Foundational Trajectory Telemetry & Operational Mappings: We investigate candidate operational correspondences connecting static algebraic-geometric properties under Singular Learning Theory (SLT) - parameter fibers, fiber volumes, and the Real Log Canonical Threshold (λ) - to dynamic differential-geometric telemetry on attention-induced Riemannian manifolds. These mappings are treated as operational hypotheses rather than established mathematical equivalences. Specifically, we propose that parameter redundancies (padding functions) may correspond operationally to the residual component of a coordinate-disentangled ANOVA hidden-state decomposition; that the macroscopic order parameter R_t may serve as an empirical probe for structural phase transitions; and that the Alpha Potential Well U(α) acts as a curvature-motivated regularizing potential against scaling symmetries. Kramers–Moyal expansions and Pawula-based diagnostics provide empirical support, but not proof, for a low-order drift–diffusion description, with higher-order diffusion terms remaining small relative to second-order variance Dʌ(4)≪(Dʌ(2)ʌ2). Extended high-resolution auditing of MyceliaLM (1.62B parameters) reveals Persistent Individual Ballistic Flow (PIBF), wherein token representations move along deterministic, drift-dominated geodesics (ρ_dir>100) while cross-token directional alignment remains near zero (c_≈0.000) and head variance spikes (σ_headʌ2≈0.730–0.760), consistent with Spontaneous Symmetry Breaking into a specialized geometric routing swarm. Across 238 snapshot audits, the relative variability of the Optimization Response Function χ_R exceeds that of directional SDE SNR ρ_dir by a factor of >148×, establishing Kinematic-Thermodynamic Decoupling as an empirical phenomenon. Furthermore, near-zero correlations between structural differentiation velocity (dσ_headʌ2/dt) and scalar loss shifts (ΔL≈0) motivate the Hidden Descent Hypothesis: internal representational specialization can proceed independently of scalar loss at the Epistemic Floor. This is the subject treated by the present paper, Beyond Scalar Loss: An Empirical Observatory for Persistent Ballistic Flow and Hidden Descent in Transformers. Closed-Loop Pre-Training Completion & Stopping Standard: To solve open-loop pre-training stopping ambiguity when scalar loss plateaus at the Epistemic Floor (L_total = L_epistemic + L_aleatoric), we introduce the Geometric Convergence Index (GCI). The GCI audits six coordinate-disentangled information-geometric criteria: inter-head fiber curvature saturation (σ²_head), differentiation velocity decay (dσ²_head/dt), attention entropy variance stability, governor work rate minimization (Ẇ_gov), Pawula Fokker–Planck stability (R_Pawula), and macroscopic order stability (R_t). The GCI defines a physics-grounded stopping standard, the Crystallization Horizon (GCI = 1.0 for 10,000 consecutive steps), that terminates pre-training the exact moment representational maturity is reached, preventing premature halt or multi-million-dollar compute waste. This is the subject of the paper, Beyond Loss Convergence: The Geometric Convergence Index and the Crystallization Horizon of Transformer Manifolds. In-Vivo Runtime Governance & Silicon Alignment: Extending forward-pass geometry monitoring to inference, we introduce the Kinematic Friction Rail and the Cache-Fused Kinematic Rail (CFKR). By monitoring depth-dependent hidden-state variance differentials via the Friction Delta (Δ = V_late - V_early), CFKR detects Geometric Flatlines (Δ → 0.00) in vivo prior to token emission, resolving the Permissive Consensus Paradox. Rather than relying on software-level additive logit bias matrices (M_ghost) and token eviction, CFKR executes ke