The evidence supports single-direction decodability for the studied task groups but challenges fixed-direction stability: the information persists while its geometric realization changes.
Abstract
Recent work has proposed that reasoning and memorization in language models can be characterized by a single representation direction, including methods that keep this direction fixed during reinforcement learning. We test two assumptions behind this view. First, are reasoning-oriented and factual-recall task groups approximately single-direction separable? Second, does the resulting geometry remain stable after GRPO? Using Qwen3-0.6B and a controlled 400-example dataset, we find that a one-dimensional projection can match a full 1024-dimensional linear probe with AUROC = 1.00 on the studied task groups. However, after GRPO, the corresponding direction is substantially reorganized: mean-direction cosine averages 0.453, probe-direction cosine 0.445, while direct representation drift reaches 0.511 at the final layer. Probe AUROC nevertheless remains 1.00. The evidence therefore supports single-direction decodability for the studied task groups but challenges fixed-direction stability: the information persists while its geometric realization changes.
This work extends contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and finds that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively.
Michal Korniak, Kamil Dybek, Benjamin Eysenbach et al.· 0 citations
Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal-aligned reinforcement inspired approach for language-model learning. Inspired by the comp...
Namkyung Yoon, Sanghong Kim, Hwangnam Kim· arXiv.org· 0 citations
Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise rel...
M. Liang, Xin-Zhao Cheng, Faizan Wajid· 0 citations
Together, these results show that composing complementary mechanisms substantially improves long-horizon memorization beyond what any individual mechanism achieves.
Zhe-Yuan Zhang, Alvin Zhang, Daniel Khashabi et al.· 1 citation
This work tracks formation step by step of a Pythia-410M model from scratch and runs attribution patching at every step, alongside probes for gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains, confirming the hypothesis that modularity tracks le...
This work presents two converging lines of evidence that linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations.