Skip to content

Author

Gleb Gerasimov

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

You Do Not Fully Utilize Transformer's Representation Capacity

Layer-Integrated Memory (LIMe) is introduced, a lightweight extension that leverages existing key-value buffers and learns per-head, per-layer routing weights to integrate representations from previous layers to improve perplexity per FLOP and yield strong gains on synthetic tasks while preserving higher value-vector entropy and token separability.

Gleb Gerasimov, Yaroslav Aksenov, Nikita Balagansky et al. · 4 citations
#machine learning Preprint Aug 2026

A Model with No Head and Many Thoughts

Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized. Experiments on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B show that Soft Latent Thinking consistently improves pass@k across all k while reducing per-step compute during chain-of-thought. Our method achieves the highest pass@32 among all soft-thinking approaches, demonstrating that effective reasoning can be carried out in continuous space without discrete token generation.

N. Koriagin, Yaroslav Aksenov, George Bredis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.