Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
Temporal Self-Distillation (TSD), a simple on-policy method that trains dLLMs for fast inference by distilling predictions across time, substantially shifts the speed--quality frontier toward the low-compute regime.
Shi-Jia Xu, Andrea Miele, Metod Jazbec et al.
· 0 citations