Rethinking Privileged Information in On-Policy Self-Distillation
It is found that stronger alignment attributable to the correct reference does not reliably coincide with a greater performance benefit from the reference, and performance gains and distributional alignment alone cannot determine how privileged reference information contributes to student learning in OPSD.