Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
CaRE-KD is proposed, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization and provides a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce.