Skip to content

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work proposes Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm that encourages each expert to operate collaboratively in the quantized model, thereby improving the overall MoE performance and reducing the dependence on the calibration dataset.

Abstract

In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. Therefore, we propose Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm. This approach encourages each expert to operate collaboratively in the quantized model, thereby 1) improving the overall MoE performance and 2) reducing the dependence on the calibration dataset. Since uniformly adjusting each expert's performance facilitates robustness and stability of the MoE model, the proposed MoE quantization method can generalize more consistently across different calibration datasets. Our code is available at: https://github.com/mmai-laboratory/Colla_Q

View source

Similar papers

#machine learning Preprint Sep 2026

Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution

Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...

Ning-Feng Yang, T. Aamodt · 0 citations
#artificial intelligence Preprint Aug 2026

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

A novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs and a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strateg...

G. Lee, Ji Liu, Juncheng Jia et al. · 0 citations

OJBKQ : Objective-Joint Babai–Klein-Based Quantization ∗

OJBKQ is proposed, a layer-wise PTQ method that formulates weight quantization as a joint optimization problem over activations and weights, yielding a multiple-right-hand-side box-constrained integer least squares (BILS) problem per layer.

Xin-Yu Wang, Zi-Yu Zhao, Peng Lu et al. · 0 citations
Book Open access Aug 2026

Quantized Model Soup Shake-Up: Weight Perturbation for Enhanced Ensemble Diversity

Model soup, averaging the weights of multiple fine-tuned models, delivers ensemble-level accuracy at single-model inference cost, but its success requires both linear mode connectivity (LMC) and sufficient diversity among candidates. We study these two requirements under quantization-aware training (QAT). First, we sho...

Jinwook Chung, Sungyeop Jung, Weronika Czorapinska et al. · 0 citations
2026

Distribution-aware Low-bitwidth Quantization for Large Language Models

A comprehensive PTQ framework is presented that addresses the problem of compressing LLM weights through three core innovations: a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, a learnable codebook optimization mechanism employing noise substitution...

Bao Tan Duy Huynh, Takashi Tsunakawa, Masafumi Nishida · 0 citations
#machine learning Preprint Sep 2026

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...

Jun-Hao Hu, S. Ramachandran · 1 citation

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.