Skip to content

Author

Zun-Hai Su

We have 3 of 7 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their di...

Zun-Hai Su, Yuxuan Sun, Jian-Chao Tan et al. · 0 citations
Preprint Aug 2026

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem, is introduced, and a layer-adaptive protection mechanism that restricts the merging of high-traffic experts based on their routing exposure is introduced.

Lei Xin, Bin Gu, Peize Li et al. · 3 citations
#small language model Preprint Aug 2026

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

DAMP uses both quantization-error energy and decay-based persistence to identify high-risk channels during offline calibration and stores these channels at higher precision and the remainder in INT8, the first to study post-training quantization of recurrent states in GDN and KDA based language models.

Tao Zhang, Jian-Chao Tan, Ping-Wei Sun et al. · 4 citations · ⚡2

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.