Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
A unified latency decomposition framework for dLLMs is introduced to disentangle factors and analyze their impact on inference speed in real deployments, and categorize acceleration techniques along three axes covering algorithmic innovations, architectural and system optimizations, and inference-time scaling.