Skip to content

Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head

Sep 2026 · 0 citations · 52 references
Computer Science

TL;DR

NAMOH, an architecture-native sparse attention mechanism that activates only its assigned tokens and performs causal attention within this subsequence, is introduced, and it is hoped this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.

Abstract

Scaling attention parameters can improve language model quality, but retaining full token histories makes additional heads costly at long contexts. Furthermore, since attention retrieves and combines contextual information, parameter scaling should also support longer contexts. We therefore ask whether attention parameter scaling can directly enable efficient and effective context scaling. We introduce NAMOH, an architecture-native sparse attention mechanism that activates $K$ of $H$ heads per token. Each head retains only its assigned tokens and performs causal attention within this subsequence. Head selection thus jointly determines active parameters and available context without scanning the full history. Under balanced assignments, increasing $H$ at fixed $K$ shortens head histories and reduces per-token key-value (KV) access without increasing total KV storage. We further support head-relative rotary position embeddings to shorten positional spans within routed subsequences, aiming to mitigate position-induced attention noise. Experiments show that NAMOH can outperform fully activated models with the same total parameters, while enabling more efficient long-context inference than smaller dense models with matched active parameter counts. It remains compatible with GQA and existing sparse attention mechanisms. We hope this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.

View source

Similar papers

#machine learning Preprint Aug 2026

LoGo: Token-Level Dynamic Local-Global Attention

LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.

Yu-Qi Pan, Zheng Li, Bo-Hao Tang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

RBS-Attention is proposed, a training-free sparse-prefill method with two complementary selection branches that controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution and supports radius-adaptive dual-branch selection as an effective approach to long-context prefill.

Chu-Xu Song, Jiu-Qi Wei, Zhen-Can Peng · 0 citations
#machine learning Preprint Sep 2026

When Can Attention Heads Be Statically Defined?

Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance a...

Weixian Waylon Li, Yin-Tao Tai, Marcio Fonseca et al. · 0 citations
#machine learning Preprint Sep 2026

Retrieval Capacity of Self-Attention Under Competition

How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the s...

Timur Mudarisov, M. Burtsev, Radu State · 0 citations
#artificial intelligence Preprint Aug 2026

DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

Decay-Aware State Compression (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout to integrate efficiently with tensor-parallel inference engines.

Yanzhi Yu, Ping-Wei Sun, Jian-Chao Tan et al. · 2 citations
#machine learning Preprint Sep 2026

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

This work introduces Elastic Threshold Attention, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding, and introduces an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds, cutting attention compute by an...

Themistoklis Haris, Henry Li, Maryam Karimzadehgan · 1 citation · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.