Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference
Self-Indexing Attention is proposed, a training-free framework built on a shared transform-domain sign-magnitude representation that provides a reusable token-level index for grouped prefill selection and decode retrieval, while the same representation remains compatible with external KV-cache compression without separ...