Skip to content

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, the framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.

Abstract

Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. We trace this gap to quantization-amplified exposure bias: QAD trains on fixed corpus prefixes, while quantization-induced deviations compound along the model's own autoregressive trajectories. To address this mismatch, we introduce an on-policy distillation (OPD) stage that places teacher supervision where the quantized model actually goes. Starting from a QAD checkpoint, the student generates through the quantized forward path used at deployment and receives feedback from a frozen full-precision teacher on its own prefixes, combining dense token-level guidance with task-verifier rewards. Across four models at 2.79 and 1.88 effective bits, OPD raises average BF16 performance retention from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval while preserving short-form performance, with reasoning gains substantially exceeding those of continued teacher-forced QAD in matched-budget comparisons. By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, our framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by its...

Wen-Jun Wang, He-Ping Li, Yang-Gan Gu et al. · 0 citations
#machine learning Preprint Sep 2026

Activation-Conditioned Self-Distillation

On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distilla...

Zhe-Xi Lu, Subhajit Chaudhury, Tejaswini Pedapati et al. · 0 citations
#machine learning Preprint Oct 2026

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optim...

Karn Tiwari, V. Chordia, P. PrathoshA · 0 citations
#machine learning Preprint Sep 2026

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation, shows that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

Safaeid Hossain Arib, Rabeya Akter, I. N. Swapnil et al. · 0 citations
Preprint Sep 2026

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We fi...

Zi-Zhuo Lin, Quan-Ling Liu, Yi Yang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.