Preprint
Aug 2026
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
AdaMX (Adaptive Microscaling), a heterogeneity-aware format and accelerator that removes 83% of the MXFP4 accuracy loss on commonsense and 82% on MMLU, and 43% and 27% of the NVFP4 loss across LLMs from 3B to 70B.
Junyi Luo, Xin Jiang, Tai-Hao Wen et al.
· 0 citations