BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization
BitsMoE is proposed, a cost-aware mixed-precision quantization framework built on two complementary techniques that separates expert weights into a shared basis and expert-specific spectral components, defining structural quantization units while exploiting cross-expert redundancy.
Jiayu Zhao, Zi-Han Teng, Min-Hao Fan et al.
· 0 citations