Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning
By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, the framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.