Skip to content
Preprint

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

The results show that abstract skills provide dense supervision where group-relative rewards become uninformative.

Abstract

Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. We propose SKALD (Skill-Anchored Latent Distillation), an on-policy self-distillation framework that uses two context views of the same Qwen3-Base model: a question-only student and a teacher conditioned on an abstract, explicit-answer-filtered skill card. The student is trained on its own prefixes, transferring the skill-induced advantage into shared parameters without privileged input at test time. To stabilize context-induced distribution mismatch, SKALD employs an annealed exponentially tilted objective that downweights teacher-preferred tokens with very low student likelihood; as the tilt vanishes, it converges to teacher cross-entropy and recovers the forward-KL student gradient. An empirical gate activates distillation only when verified rollouts estimate a positive teacher advantage. Across five held-out mathematics benchmarks, SKALD improves overall avg@8 over GRPO by +2.46, +4.85, and +12.01 at 0.6B, 1.7B, and 4B, respectively. At 1.7B, zero-variance-only distillation recovers 84.7% of the full gain, while SKALD remains +4.06 above FLOP-matched GRPO and exceeds contextual skill exposure by +3.77. These results show that abstract skills provide dense supervision where group-relative rewards become uninformative.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance, and demonstrates that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance.

Zhu Zhang, Ji-Xun Wang, Xiao-An Xu et al. · 1 citation
Preprint Aug 2026

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation

This work reproduces SDPO's reported gains in its easy setting, then applies the identical setup to difficult tasks and finds that it does not teach anything, and explains this failure through a single causal chain from the loss to the model it produces.

Sarthak Harne, Chinmay Karkar, Yash Pandya et al. · 3 citations
Review Aug 2026

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

This review treats collapse as a symptom governed by three levers: where the signal is applied, that is, how tokens are weighted; what the teacher is shown, that is, the nature of the privileged information; and when the signal changes, that is, the teacher's dynamics and the decay of guidance.

J. Robert, Raheel Qader · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.