Skip to content
Preprint

Measuring Structured Predictability in Neural Training Dynamics: A Cross-Regime Study

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

This work combines three complementary probe families, displacement-direction, subspace-residual, and predictor-based probes, with convention-aware, null-calibrated group-level readouts, and applies them to multi-pass vision training on CIFAR and public Pythia pretraining checkpoints, indicating that these measurements capture intrinsic trajectory structure, while probe differences distinguish complementary forms of temporal organization.

Abstract

Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. We study short-horizon predictability as a measure of temporal redundancy: where, when, and under which training conditions recent updates contain information about near-future parameter motion. We combine three complementary probe families, displacement-direction, subspace-residual, and predictor-based probes, with convention-aware, null-calibrated group-level readouts, and apply them to multi-pass vision training on CIFAR and public Pythia pretraining checkpoints. Across both regimes, vector-like tensors such as normalization parameters and biases (auxiliary parameters) exhibit simpler short-horizon dynamics than matrix-like feature-transforming weights (bulk parameters), whose predictable behavior concentrates in localized, time-varying pockets. Agreement within and across probe families, and with independent trajectory diagnostics, indicates that these measurements capture intrinsic trajectory structure, while probe differences distinguish complementary forms of temporal organization. Controlled CIFAR comparisons further show that architecture and training recipe systematically modulate the measured structure. A Pythia-70M case study further exposes a sequence of role-, depth-, and scale-dependent events, including bulk ESA falling below the random sign-agreement level and the emergence and redistribution of predictable qkv pockets across layers. These results position short-horizon predictability as a retrospective, parameter-resolved diagnostic of training dynamics.

View source

Similar papers

Open access Aug 2026

Dynamic Regime Maps of Neural Network Training Under the Adam Optimizer: An Observable-Based Empirical Analysis

Adaptive optimization algorithms are fundamental to modern deep learning; however, the global organization of neural-network training regimes induced by optimizer hyperparameters remains insufficiently understood. In particular, the influence of the Adam moment coefficients on the stability and qualitative behavior of the learning process has not been systematically investigated through parameter-space regime mapping. In this work, we introduce an observable-based empirical regime-mapping framework for analysing neural-network training under the Adam optimizer in the two-dimensional hyperparameter space defined by the exponential decay coefficients of the first and second gradient moments, (β 1 , β 2 ). Neural-network optimization is treated as an iterative parameter-update process evolving in a high-dimensional parameter space, while its behavior is characterized through low-dimensional observable fields derived from neuron-wise training-error dynamics. Rather than relying on a single observable, the proposed framework combines seven complementary empirical descriptors that characterize regime organization, temporal stability, alignment, anisotropy, and training evolution. Experiments are performed on datasets of increasing complexity, including printed-digit patterns, Fashion-MNIST, CIFAR-10, and CIFAR-100, using multilayer neural-network architectures of varying width and depth. The resulting regime maps reveal fragmented hyperparameter landscapes containing stable, oscillatory, slow-learning, and irregular observable regimes. Increasing dataset complexity and network capacity is generally associated with smaller coherent stable regions and increased sensitivity to the Adam moment coefficients. The geometric complexity of the regime boundaries is quantified using box-counting analysis. The estimated dimensions approach D ≈ 1.9 for several investigated configurations, indicating highly irregular and nearly space-filling boundaries at the available numerical resolution. These values are interpreted as empirical measures of boundary complexity rather than as evidence of exact mathematical fractality. Additional robustness experiments performed using 300, 500, and 1000 optimizer steps demonstrate that the large-scale organization of all seven observable fields remains largely preserved, whereas the principal changes are concentrated near transition boundaries. This persistence indicates that the detected regime structures are reproducible and are not solely artifacts of short optimization histories. The proposed methodology provides an empirical computational framework for visualizing and comparing optimization regimes in the Adam hyperparameter space. It facilitates the identification of comparatively stable hyperparameter regions and offers a complementary observable-based perspective on the complex behavior of adaptive neural-network optimization.

S. Sveleba, I. Katerynchuk, I. Kunyo et al. · 0 citations
Preprint Aug 2026

Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW

This work forms AdamW as a finite-horizon input--state--output (ISO) system whose state contains the model parameters and first- and second-moment estimates, and derives an exact multistep error decomposition and establishes first-order finite-horizon accuracy under local smoothness and controlled activation switching.

Kangning Liu, Suyan Li · 0 citations
Open access Aug 2026

Task-Parametrized dynamics: Representation of time and decisions in recurrent neural networks.

Training RNNs on delayed decision-making tasks with progressively increasing temporal demands shows that temporal and decision-related computations can emerge through multiple dynamical regimes, while maintaining structured low-dimensional representations and comparable behavioural performance, mirroring biological principles of degeneracy and functional redundancy.

Cecilia Jarne, Ryeongkyung Yoon, Tahra L. Eissa et al. · 0 citations
Jul 2026

Neural operator discovery from heterogeneous trajectories

A factorized latent-conditioning formulation is introduced that jointly learns a neural operator and a low-dimensional latent representation through factorized prediction, trajectory-decoupled sampling, and dimension selection that enables generalization to previously unseen system instances.

Zi-Tuo Chen, Qiaofeng Li, Jia-Xin Hu et al. · 1 citation
#machine learning Preprint Aug 2026

Learning Generalizable Reconstruction of High-Dimensional Neural Dynamics

Accurate reconstruction of long-duration neural recordings is challenging because local field potentials (LFPs) are high-resolution, multichannel, transient, and variable across subjects. We present PCA-DMD, a scalable operator-theoretic framework that segments LFP recordings into overlapping windows, projects them into a compact PCA space, learns linear Koopman evolution in the latent space, and reconstructs continuous signals through inverse projection and overlap-add aggregation. On 200,000-sample hippocampal recordings, PCA-DMD outperformed Classical DMD, SpDMD, MrDMD, and HODMD, achieving KLD=0.0761 and HD=0.0847. In all-pair cross-subject zero-shot generalization at 300,000 samples, correlations were 0.9504-0.9800, with HD=0.0010-0.0072 and KLD=0.0005-0.0022, without target-subject fine-tuning. Out-of-sample temporal prediction showed close one-step agreement on temporally held-out LFP segments across the unseen interval and multiple channels. Scalability analysis from 400,000 to 900,000 samples showed stable zero-shot reconstruction, with mean correlation remaining about 0.965-0.968 while computational cost increased predictably. External validation on an independent 93-channel Allen Neuropixels recording yielded mean and median channel-wise correlations of 0.7427 and 0.7990, respectively. Koopman spectral and mode analyses revealed dominant eigenvalues concentrated near the unit circle. PCA-DMD therefore provides an interpretable, generalizable, and computationally scalable framework for reconstructing high-dimensional neural dynamics.

Anima Kujur, Z. Monfared · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.