Skip to content

Author

Florent Krzakala

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

High-Dimensional Learning Dynamics of Attention-Indexed Models

Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix $S\in\mathbb{R}^{d\times d}$ can remain trapped in an uninformative state. Tied attention ($S=WW^\top$) induces an automatic symmetry-breaking mechanism and yields weak recovery in $\Theta(d^2\log d)$ samples. For untied attention, $S=UV^\top$, we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the $\Theta(d^2\log d)$ scale occurs when the state selected by the fast dynamics breaks the initial symmetry.

Yizhou Xu, M. Sagitova, Lenka Zdeborová et al. · 0 citations
Preprint Aug 2026

Spectral phase transitions in Gaussian multi-index models

Recovering a low-dimensional latent subspace from nonlinear observations of Gaussian covariates in high dimensions is a fundamental problem in feature learning. Here, we consider Gaussian multi-index models in which the covariates $\boldsymbol{x}_i \stackrel{\mathrm{i.i.d.}}{\sim} \mathcal{N}(0,\boldsymbol{I}_d)$ and the responses $\boldsymbol{y}_i$ depend on $\boldsymbol{x}_i$ only through its projection onto an unknown $r$-dimensional subspace. Earlier work based on approximate message passing (AMP) identified a sharp threshold for weak recovery [Troiani et al., 2025], raising the question of whether it can be attained, without side information, by a spectral method. We answer this affirmatively and develop a general random matrix theory for matrix-valued spectral estimators of the form \[\boldsymbol{D}_n=\frac{1}{n}\sum_{i=1}^n\boldsymbol{T}(\boldsymbol{y}_i)\otimes\boldsymbol{x}_i\boldsymbol{x}_i^\top,\] where $\boldsymbol{T}$ is an arbitrary bounded symmetric matrix-valued preprocessing map of fixed dimension. As $n,d \to \infty$ with $n/d\to\alpha$, we prove that the empirical spectral measure of $\boldsymbol{D}_n$ converges almost surely to a deterministic compactly supported distribution characterized by a matrix-valued self-consistent equation. We then establish a spectral phase transition for the largest eigenvalue: below threshold it sticks to the bulk edge, while above threshold an outlier emerges. We characterize the outlier location through a finite-dimensional deterministic equation and show that the associated spectral estimator achieves weak recovery of the latent subspace. Finally, we prove that the AMP-derived preprocessing of [Defilippis et al., 2025] is optimal among all bounded matrix-valued preprocessing maps of any fixed dimension. Its transition coincides with the AMP weak-recovery threshold, proving the general spectral conjecture of [Defilippis et al., 2025].

Florent Krzakala, Pierre Mergny, Vanessa Piccolo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.