Skip to content

Efficient Reasoning via Constrained Optimization in Latent Space

Sep 2026 · 0 citations · 47 references
Computer Science Mathematics

TL;DR

This paper proposes a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance, and investigates the latent representations and proposes a novel training-free framework.

Abstract

Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.

View source

Similar papers

Preprint Aug 2026

ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

This work proposes ChainPrune, a novel reasoning path semantic structural optimization method to efficiently and controllably synthesize self-generated high-quality training data and incorporates a DPO-based preference learning method combined with supervised loss, effectively mitigating false reward suppression.

Weihang Pan, Zhengxu Yu, Yuxiang Zhang et al. · 1 citation
Conference Open access Sep 2026

Sprint or Delve: A Distribution-Aware Approach to Efficient Reasoning

The Powered Length Penalty (PLP) is proposed, an adaptive regularizer that penalizes redundancy in short sequences while gradually reducing penalties for longer sequences, preserving deep reasoning.

Ze-Hui Ling, De-Shu Chen, Hong-Wei Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

Efficient Reasoning Exploration via State-Conditioned Latent Steering with Progress Guidance

SPS is a training-free latent steering framework that constructs a state-conditioned Direction Bank containing multiple progress-guided steering vectors for different prefix-state regions and applies it at high-uncertainty transitions to guide the next reasoning step toward meaningful progress.

Hengyuan Zhang, Chenming Shang, Zun-Hai Su et al. · 0 citations
#machine learning Preprint Sep 2026

Assessing Adversarial Robustness of Latent Reasoning Models

Large language models increasingly rely on long chain-of-thought (CoT) trajectories for complex reasoning, but autoregressive generation brings substantial memory and inference costs. Latent reasoning models (LRMs) offer a more efficient alternative by compressing intermediate reasoning into a small number of continuou...

Shao-Long Chen, Ang Li, Ming-Jie Li et al. · 0 citations
#machine learning Preprint Oct 2026

What Matters for Latent Reasoning with Flow Matching

Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reason...

Yassine Ouali, Adrian Bulat, G. Tzimiropoulos · 0 citations
#natural language process... Preprint Oct 2026

Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a la...

Ming-Qing Yuan, Xiao-Bo Liang, Jun-Wei Yang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.