Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Flash-dLLM is introduced, a training-free inference acceleration framework for fast and memory-efficient dLLMs that proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model.
Quan Nguyen-Tri, Mukul Ranjan, Zhi-Qiang Shen
· 0 citations