Skip to content

Reasoning Fine-Tuning Induces Persistent Latent Policy States

Jul 2026 · arXiv.org · Vol abs/2607.18532 · 0 citations · 39 references
Computer Science

TL;DR

This work modeling Chain-of-Thought reasoning as a switching dynamical system (SDS) suggests that reasoning fine-tuning globally reorganizes latent dynamics, offering a new lens for mechanistic analysis and process-level control of reasoning models.

Abstract

Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood. It is unclear whether reasoning fine-tuning improves local token-level competence or globally reorganizes how models structure inference over time. We address this question by modeling Chain-of-Thought reasoning as a switching dynamical system (SDS), in which internal representations evolve under discrete latent policy states. Our framework combines time-aware contrastive representation learning with discrete regime discovery to recover latent policies from activation trajectories. Across four benchmarks and model scales from 1.5B to 32B parameters, reasoning-fine-tuned models exhibit richer latent-policy organization than their base counterparts, characterized by more differentiated transition structure and model-dependent changes in state utilization, persistence, and mixing. The recovered regimes exhibit functional specialization aligned with distinct reasoning stages, and extensive controls confirm that their structure is not explained by correctness, representation learning, or modeling priors, but depends on the coherent temporal organization of reasoning trajectories. Causal interventions further show that the regimes are functionally meaningful: state-swap ablations reduce one-step predictive fit, while transplanting reasoning dynamics into base models improves performance on challenging reasoning problems. Finally, SDS-guided pruning of failure-prone reasoning prefixes outperforms self-consistency in 11 of 12 model-dataset settings, with gains of up to 12.5 percentage points. Together, our results suggest that reasoning fine-tuning globally reorganizes latent dynamics, offering a new lens for mechanistic analysis and process-level control of reasoning models.

View source

Similar papers

Jul 2026

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

This work treats each reasoning trace as a sequence of latent states rather than an unstructured texts, and investigates whether inference time interventions can provide fine-grained control over the self-looping reasoning process.

Sheldon Yu, Tong Yu, Xunyi Jiang et al. · 0 citations
Jul 2026

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Surrogate Latent Policy Optimization (SLPO) is introduced to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy.

Runyang You, Zhiyuan Liu, Yongqi Li et al. · 1 citation
Jul 2026

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

LatentRM is a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards through on-policy optimization of the latent reasoning space end-to-end.

Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang et al. · 0 citations
Jul 2026

In-Context Learning as Implicit Policy Gradient

It is shown that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization, and an exact upper bound on the distribution shift induced by a bounded attention update is derived, yielding a trust-region-like analogy to KL-constrained policy optimization.

Masahiro Kaneko, Timothy Baldwin · 0 citations
Preprint Aug 2026

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

GradCuit (gradient through circuit), which inserts optimizable latent states at a selected Transformer layer between the hidden representations of the prompt and the generated continuation, opens a new axis of robust and interpretable test-time scaling, where LLMs adapt how they reason rather than merely regenerate, sample, or rerank outputs.

Zhaoxin Yu, Qianli Shen, Hengli Li et al. · 0 citations
Jul 2026

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

This work presents two converging lines of evidence that linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations.

Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.