Skip to content

Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold

Jul 2026 · arXiv.org · Vol abs/2607.27770 · 1 citation · 31 references
Computer Science

TL;DR

Strong students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union, in an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation.

Abstract

A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore be viewed as local probes of a multi-basin reasoning solution manifold, rather than as globally reliable supervisors. Based on this view, we propose an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation. In the expansion stage, Residual Group Relative Policy Optimization (RGRPO) trains a sequence of teachers from a common initialization and redirects each later round toward examples not yet covered by the accumulated teacher union. In the compression stage, reliability-gated Teacher-Union On-policy Distillation (TU-OPD) lets the student learn from its own response prefixes. For each example, only reliable teachers contribute, and their sampled-token OPD losses are weighted by their per-example quality. We further introduce Consensus-Residual Decomposition, which preserves a winner teacher's excess token preferences over its reliable peers, preventing specialist behavior from being suppressed during teacher aggregation. Experiments on mathematical reasoning, code generation, and instruction following show that the resulting Qwen3-1.7B student consistently outperforms the strongest individual teacher across all three domains, yielding relative improvements of 2.0%, 8.3%, and 6.9%, respectively, while retaining single-model inference. These results establish a simple but powerful principle: stronger students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union.

View source

Similar papers

Preprint Aug 2026

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

Persistent Consistency Self-Distillation (PCSD) is proposed, which derives token-level distillation weights from the local persistence of teacher-favoring signals, and combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support.

Chunji Lv, Yangguang Wei, Junlin Liu et al. · 0 citations
Jul 2026

Weak-to-Strong Generalization via Direct On-Policy Distillation

Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.

Shiyuan Feng, Huan Gao, Haohan Chi et al. · 8 citations · ⚡1
Jul 2026

Reward-Gated On-Policy Distillation

RG-OPD bridges sparse verifier rewards and dense teacher logits, preserving token-level supervision while filtering misleading teacher signals, and produces stronger distilled students, outperforming both vanilla reverse-KL distillation and the recent TSD-KD baseline.

Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi et al. · 4 citations · ⚡1
Preprint Aug 2026

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

This work proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward.

Chen Yang, Haiyuan Wan, Rengrong Xiong et al. · 0 citations
Preprint Jul 2026

Geometric Self-Distillation for Reasoning Generalization

GeoSD, a geometric self-distillation objective that treats this drift as movement in the student's predictive behavior and counters it in two complementary ways, preserves the in-distribution gains of self-distillation while improving average OOD accuracy.

Josip Jukic, Ivan Titov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.