Commutator memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training, which identifies which came from which order in 92% of cases across four LLMs (chance 50%).
Abstract
Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leaves a parametric training-history memory: a weight component that flips sign when the two sources are swapped, is localized in output space, changes the held-out loss gap between the two orders under targeted interventions, and reveals which trained model came from which order. For one small SGD step of size $\eta$ on each of sources $A$ and $B$, the weight difference $\theta_{AB}-\theta_{BA}$ is, to leading order, $\eta^2 b_{AB}$, where $b_{AB}=H_Bg_A-H_Ag_B$ is the Lie bracket of the two gradient fields at the base model. We define commutator memory by projecting the bracket through the logits into one score per vocabulary token; the scores sum to the bracket's prediction of the gap. The scores are localized: on three models, the same readout of the measured $\theta_{AB}-\theta_{BA}$, or of a bracket from disjoint batches, shares 82-99% of the original top-20 tokens, versus 35-49% for norm-matched random directions. They are causally actionable: in Qwen-3-4B SFT, downweighting the ten tokens with the largest predicted share of the gap closes a median 32% of the measured gap, while frequency-matched tokens with near-zero scores have almost no effect. The weights themselves carry the component: projecting the difference between the two trained models onto $b_{AB}$ identifies which came from which order in 92% of cases across four LLMs (chance 50%). Controlled tests also cover matched-batch DPO, a frozen-rollout GRPO-style objective, and an AdamW endpoint check. The memory is defined per source pair, not per example, and its projection on $b_{AB}$ decays with further training.
The RunningTensor is introduced, which generalizes this memory to an order-o tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries, and improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide usef...
We introduce a simple architectural modification to decoder-only transformers: a persistent recurrent state that observes hidden representations via cross-attention, updates itself through a GRU, and modulates subsequent processing via gated addition. Inserted between the lower and upper halves of a 6-layer transformer...
Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectral norm is not a faithful measure of functional change. Softmax removes shared logit shifts, whereas the spectral norm can assign arbitraril...
Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno et al.· 0 citations
This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.
Yi-Fan Zhang, Steve Ta, Jasper Zhang et al.· 0 citations
Controls show that repeated-token failures are confounded by the number of independently sampled symbols, that adding never-target output classes does not impair learning, and that projection diagnostics do not reliably predict progress in the authors' runs.
Can gradient-based training learn the rank needed to store and compose associations in a matrix memory? In our earlier study, we used a matrix-augmented reasoner on a task that admits a rank-1 solution, leaving this question open. We train matrix memories on $K$ fresh key-value bindings whose exact linear recovery requ...