This work proposes Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm that encourages each expert to operate collaboratively in the quantized model, thereby improving the overall MoE performance and reducing the dependence on the calibration dataset.
Abstract
In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. Therefore, we propose Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm. This approach encourages each expert to operate collaboratively in the quantized model, thereby 1) improving the overall MoE performance and 2) reducing the dependence on the calibration dataset. Since uniformly adjusting each expert's performance facilitates robustness and stability of the MoE model, the proposed MoE quantization method can generalize more consistently across different calibration datasets. Our code is available at: https://github.com/mmai-laboratory/Colla_Q
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...
A novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs and a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strateg...
OJBKQ is proposed, a layer-wise PTQ method that formulates weight quantization as a joint optimization problem over activations and weights, yielding a multiple-right-hand-side box-constrained integer least squares (BILS) problem per layer.
Xin-Yu Wang, Zi-Yu Zhao, Peng Lu et al.· 0 citations
Model soup, averaging the weights of multiple fine-tuned models, delivers ensemble-level accuracy at single-model inference cost, but its success requires both linear mode connectivity (LMC) and sufficient diversity among candidates. We study these two requirements under quantization-aware training (QAT). First, we sho...
Jinwook Chung, Sungyeop Jung, Weronika Czorapinska et al.· Proceedings of the 32nd ACM...· 0 citations
A comprehensive PTQ framework is presented that addresses the problem of compressing LLM weights through three core innovations: a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, a learnable codebook optimization mechanism employing noise substitution...
Bao Tan Duy Huynh, Takashi Tsunakawa, Masafumi Nishida· International Conference on...· 0 citations
Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...
Jun-Hao Hu, S. Ramachandran· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.