Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their di...
Zun-Hai Su, Yuxuan Sun, Jian-Chao Tan et al.· 0 citations
UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem, is introduced, and a layer-adaptive protection mechanism that restricts the merging of high-traffic experts based on their routing exposure is introduced.
DAMP uses both quantization-error energy and decay-based persistence to identify high-risk channels during offline calibration and stores these channels at higher precision and the remainder in INT8, the first to study post-training quantization of recurrent states in GDN and KDA based language models.
Tao Zhang, Jian-Chao Tan, Ping-Wei Sun et al.· 4 citations· ⚡2
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.