OISD: On-Policy Internal Self-Distillation of Language Models
The OISD framework is proposed, which improves reasoning by transferring on-policy predictive signals from the final layer to intermediate representations and employs signed advantage-weighted Jensen--Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy.