We introduce Grouped-Query Latent Sparse Attention (GQLSA), a novel attention mechanism that integrates latent compression, grouped-query sharing, and block-sparse attention into a unified hardware-native architecture. GQLSA achieves linear O(T) complexity in sequence length, a 3.8x speedup over multi-head attention at sequence length 4096, a 2.2x reduction in peak memory consumption, and a 16x smaller KV cache—while maintaining language modeling quality on par with dense baselines (perplexity 13.06 versus 13.12 for MHA on WikiText-2). We present the complete mathematical formulation, algorithmic pseudocode, causal correctness verification, and comprehensive empirical evaluation across speed, memory, and quality dimensions.
Fardin Sabid· Zenodo (CERN European Organi...· 0 citations
We introduce Grouped-Query Latent Sparse Attention (GQLSA), a novel attention mechanism that integrates latent compression, grouped-query sharing, and block-sparse attention into a unified hardware-native architecture. GQLSA achieves linear O(T) complexity in sequence length, a 3.8x speedup over multi-head attention at sequence length 4096, a 2.2x reduction in peak memory consumption, and a 16x smaller KV cache—while maintaining language modeling quality on par with dense baselines (perplexity 13.06 versus 13.12 for MHA on WikiText-2). We present the complete mathematical formulation, algorithmic pseudocode, causal correctness verification, and comprehensive empirical evaluation across speed, memory, and quality dimensions.
Fardin Sabid· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.