Skip to content

Category

neuroscience

134 papers

#artificial intelligence Preprint Aug 2025

Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers

PH-attention is introduced, a Pearce--Hall-inspired rule with an explicit feature-indexed associability state that yields cue-dependent learning rates and is absent from the token-computed gates the authors compare, and predicts a dissociation that survives training on generic in-context association.

Mu Qiao · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.