2026· IEEE Transactions on Audio, Speech, and Language Processing· Vol 34, pp. 4331-4347· 0 citations· 65 references
TL;DR
Evaluations show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.
Abstract
Model quantization facilitates Automatic Speech Recognition (ASR) deployment on resource-constrained devices, yet existing methods suffer from severe accuracy loss below 4 bits, redundant storage for multiple precisions, and limited applicability across training paradigms. We address these challenges with two methods that share a hierarchical multi-precision design. When full retraining is feasible, Quantization-Aware Co-Training (QACT) jointly optimizes weight-shared 2-bit and 1-bit sub-networks. Its 2-bit Conformer achieves 12.86% WER on Switchboard, statistically equivalent to the FP32 baseline by MAPSSWE at $\alpha {=}0.05$, while supporting both precisions in one model. For foundation models where retraining is prohibitive, Codebook-Shared Vector Quantization (CSVQ) hierarchically optimizes shared codebooks at the post-training stage. CSVQ substantially improves over scalar Post-Training Quantization on HuBERT and Whisper and supports 2/3/4-bit inference without additional multi-precision storage. Evaluations across supervised Conformer, self-supervised Wav2Vec2 and HuBERT, and weakly-supervised Whisper show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.1
A comprehensive PTQ framework is presented that addresses the problem of compressing LLM weights through three core innovations: a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, a learnable codebook optimization mechanism employing noise substitution...
Bao Tan Duy Huynh, Takashi Tsunakawa, Masafumi Nishida· International Conference on...· 0 citations
Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit...
This novel progressive quantization framework combines multi-bit assistive teacher models with self-knowledge distillation to stabilize BNN training and integrates matching structured pruning with an asymmetric Binary Weight Network scaling factor, thereby reducing quantization errors while maintaining hardware efficie...
Jie Xu, Wonjun Hwang, Hyunsouk Cho· Multimedia tools and applica...· 0 citations
Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.
All for 1-Bit (AF1) is proposed, a genuine 1-bit PTQ framework for LLMs that consistently outperforms existing binarization-based PTQ methods in perplexity and zero-shot accuracy, providing a practical path toward deployable genuine 1-bit compression for LLMs.
Zhi-Xiong Zhao, Zu-Kang Xu, Guang-Yu Sun et al.· 1 citation
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...
Ning-Feng Yang, T. Aamodt· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.