#artificial intelligence
Dec 2025
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
Kascade is a training-free sparse attention method that leverages known observations such as 1) post-softmax attention is intrinsically sparse, and 2) the identity of high-weight keys is stable across nearby layers to achieve high accuracy on long-context LLM inference.
Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra et al.
· arXiv.org · 10 citations