Skip to content

Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

Commutator memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training, which identifies which came from which order in 92% of cases across four LLMs (chance 50%).

Abstract

Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $\eta$ on each of sources $A$ and $B$, the weight difference $\theta_{AB}-\theta_{BA}$ is, to leading order, $\eta^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $\theta_{AB}-\theta_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

The RunningTensor is introduced, which generalizes this memory to an order-o tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries, and improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide usef...

Luca Herranz-Celotti, Vincent Guigue · 1 citation
#natural language process... Preprint Sep 2026

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

We introduce a simple architectural modification to decoder-only transformers: a persistent recurrent state that observes hidden representations via cross-attention, updates itself through a GRU, and modulates subsequent processing via gated addition. Inserted between the lower and upper halves of a 6-layer transformer...

E. Hering · 0 citations
Preprint Aug 2026

Toward a First-Principles Update Geometry for the Language-Model Head

Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectral norm is not a faithful measure of functional change. Softmax removes shared logit shifts, whereas the spectral norm can assign arbitraril...

Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno et al. · 0 citations
#machine learning Preprint Aug 2026

Fast Weight Attention for Continual Learning

This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

Yi-Fan Zhang, Steve Ta, Jasper Zhang et al. · 0 citations
Preprint Aug 2026

Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test

Controls show that repeated-token failures are confounded by the number of independently sampled symbols, that adding never-target output classes does not impair learning, and that projection diagnostics do not reliably predict progress in the authors' runs.

A. Murugan · 1 citation
#machine learning Preprint Sep 2026

When the Gradient Sees Rank: Provable Necessity, Causal Recruitment, and Composition in Trained Matrix Memories

Can gradient-based training learn the rank needed to store and compose associations in a matrix memory? In our earlier study, we used a matrix-augmented reasoner on a task that admits a rank-1 solution, leaving this question open. We train matrix memories on $K$ fresh key-value bindings whose exact linear recovery requ...

Samuel Larson · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.