Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation
Experiments show Influence-Directed Adaptive On-Policy Distillation (IDA-OPD), rather than relying on costly full-vocabulary Forward-KL objectives, preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability.