In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tas...
Hao-Tian Gu, Yi-Zhou Xu, L. Zdeborová· 0 citations
This work studies attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures, and reveals that attention parameterization itself can act as an architectural implicit bias.
Yi-Zhou Xu, M. Sagitova, L. Zdeborová et al.· 0 citations
It is shown that a symmetry fixes what variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero we...
Zi-Yin Liu, Yi-Zhou Xu, Tomaso A. Poggio et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.