2026· International Conference on Language Resources and Evaluation· pp. 10057-10070· 0 citations· 30 references
Computer Science
TL;DR
A comprehensive PTQ framework is presented that addresses the problem of compressing LLM weights through three core innovations: a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, a learnable codebook optimization mechanism employing noise substitution for vector quantization to enable robust gradient estimation, and a layer-grouping strategy based on statistical distribution similarity to improve parameter efficiency.
We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and t...
Daria Cherniuk, A. Rudikov, B. Kashin et al.· 0 citations
All for 1-Bit (AF1) is proposed, a genuine 1-bit PTQ framework for LLMs that consistently outperforms existing binarization-based PTQ methods in perplexity and zero-shot accuracy, providing a practical path toward deployable genuine 1-bit compression for LLMs.
Zhi-Xiong Zhao, Zu-Kang Xu, Guang-Yu Sun et al.· 1 citation
Evaluations show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.
Zhao-Qing Li, Haoning Xu, Zengrui Jin et al.· IEEE Transactions on Audio,...· 0 citations
Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.
Training Large Foundation Models (LFMs), including Large Language Models and Vision-Language Models, on massive distributed GPU clusters is increasingly bottlenecked by communication overhead. While frameworks like ZeRO++ employ static quantization to reduce communication volume, they suffer from a rigid trade-off: agg...
Hong Huang, Jia-Xun Ye, Jin-Hai Yang et al.· Proceedings of the 32nd ACM...· 0 citations
This work systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks and establishes that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
Chen-Xi Zhou, Pengfei Cao, Jin Ye et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.