Skip to content
Preprint

Beyond Fixed Directions: Adaptive Representation Analysis of Reasoning and Memorization in LLMs

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

The evidence supports single-direction decodability for the studied task groups but challenges fixed-direction stability: the information persists while its geometric realization changes.

Abstract

Recent work has proposed that reasoning and memorization in language models can be characterized by a single representation direction, including methods that keep this direction fixed during reinforcement learning. We test two assumptions behind this view. First, are reasoning-oriented and factual-recall task groups approximately single-direction separable? Second, does the resulting geometry remain stable after GRPO? Using Qwen3-0.6B and a controlled 400-example dataset, we find that a one-dimensional projection can match a full 1024-dimensional linear probe with AUROC = 1.00 on the studied task groups. However, after GRPO, the corresponding direction is substantially reorganized: mean-direction cosine averages 0.453, probe-direction cosine 0.445, while direct representation drift reaches 0.511 at the final layer. Probe AUROC nevertheless remains 1.00. The evidence therefore supports single-direction decodability for the studied task groups but challenges fixed-direction stability: the information persists while its geometric realization changes.

View source

Similar papers

#machine learning Preprint Aug 2026

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

This work extends contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and finds that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively.

Michal Korniak, Kamil Dybek, Benjamin Eysenbach et al. · 0 citations
Jul 2026

PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits

Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal-aligned reinforcement inspired approach for language-model learning. Inspired by the comp...

Namkyung Yoon, Sanghong Kim, Hwangnam Kim · 0 citations
Preprint Aug 2026

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise rel...

M. Liang, Xin-Zhao Cheng, Faizan Wajid · 0 citations
#machine learning Preprint Sep 2026

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training

This work tracks formation step by step of a Pythia-410M model from scratch and runs attribution patching at every step, alongside probes for gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains, confirming the hypothesis that modularity tracks le...

Guang-Qi Li, Yongxin Li · 0 citations
Jul 2026

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

This work presents two converging lines of evidence that linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations.

Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.