Skip to content

Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

Self-Indexing Attention is proposed, a training-free framework built on a shared transform-domain sign-magnitude representation that provides a reusable token-level index for grouped prefill selection and decode retrieval, while the same representation remains compatible with external KV-cache compression without separate indexer metadata.

Abstract

Sparse long-context inference requires efficient token retrieval in both prefill and decode. Existing methods often use different retrieval strategies for the two stages, preventing one retrieval representation from being reused throughout inference. We propose Self-Indexing Attention, a training-free framework built on a shared transform-domain sign-magnitude representation. The key signs provide a reusable token-level index for grouped prefill selection and decode retrieval, while the same representation remains compatible with external KV-cache compression without separate indexer metadata. This 1-bit index enables efficient retrieval through bitwise operations widely supported by modern accelerators. At 5% attention density, Self-Indexing Attention remains close to dense attention on LongBench and RULER and achieves up to 6.1x prefill and 10.3x decode attention-operator speedups. Experiments with TurboQuant and DeepSeekV4-Flash further demonstrate compatibility with low-bit KV-cache compression and pretrained sparse-attention indexers.

View source

Similar papers

#machine learning Preprint Oct 2026

OVAL: Output-Aware Local Page Bases for KV Cache Retrieval

Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query. Existing retrieval methods are designed to estimate attent...

Ashkan Shahbazi, Chayne Thrash, Soheil Kolouri · 0 citations
#artificial intelligence Preprint Sep 2026

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

This work introduces Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization and strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.

Pei-Chun Hua, Yun-Ming Xiao · 1 citation
#machine learning Preprint Sep 2026

Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference

The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval capability varies sub...

Xian-Peng Shang, Can-Bin Huang, Jiang Li et al. · 0 citations
#natural language process... Preprint Aug 2026

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

This work provides a new method for fine-tuning models with sparse attention that works for any KV cache policy, runs on a moderate hardware budget, and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism).

Matthias W. Seeger, Zeyu Zhang, Vihang Patil et al. · 0 citations
#artificial intelligence Preprint Oct 2026

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a critical mismatch bet...

Lian Liu, Tian-Tian Zheng, You Huang et al. · 0 citations
Open access Sep 2026

Attention-enhanced vision transformer hashing for hybrid image retrieval

Large-scale image retrieval requires compact representations without substantially sacrificing retrieval accuracy. However, Vision Transformer Hashing (VTS) concatenates all output tokens before hash projection, resulting in a high-dimensional hashing head with considerable model and memory overhead. We replace this to...

Uyen Nguyen, Hoai Ba, Quynh Dao Thi Thuy · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.