Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning
By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, the framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.
Yuan-Teng Chen, Zhi-Lei Liu, Peisong Wang et al.
· 0 citations