HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing
On an 80B-A3B MoE model, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill computation and KV-cache storage.
Jian-Yu Wei, Yi-Zhao Gao, Qi-Hao Zhang et al.
· 0 citations