Skip to content
Open access

Toward Extremely Low-Bit and Multi-Precision Conformer and Speech Foundation Model Quantization

2026 · IEEE Transactions on Audio, Speech, and Language Processing · Vol 34, pp. 4331-4347 · 0 citations · 65 references

TL;DR

Evaluations show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.

Abstract

Model quantization facilitates Automatic Speech Recognition (ASR) deployment on resource-constrained devices, yet existing methods suffer from severe accuracy loss below 4 bits, redundant storage for multiple precisions, and limited applicability across training paradigms. We address these challenges with two methods that share a hierarchical multi-precision design. When full retraining is feasible, Quantization-Aware Co-Training (QACT) jointly optimizes weight-shared 2-bit and 1-bit sub-networks. Its 2-bit Conformer achieves 12.86% WER on Switchboard, statistically equivalent to the FP32 baseline by MAPSSWE at $\alpha {=}0.05$, while supporting both precisions in one model. For foundation models where retraining is prohibitive, Codebook-Shared Vector Quantization (CSVQ) hierarchically optimizes shared codebooks at the post-training stage. CSVQ substantially improves over scalar Post-Training Quantization on HuBERT and Whisper and supports 2/3/4-bit inference without additional multi-precision storage. Evaluations across supervised Conformer, self-supervised Wav2Vec2 and HuBERT, and weakly-supervised Whisper show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.1

Read PDF

Similar papers

2026

Distribution-aware Low-bitwidth Quantization for Large Language Models

A comprehensive PTQ framework is presented that addresses the problem of compressing LLM weights through three core innovations: a calibration process guided by Kullback-Leibler divergence minimization to preserve the original weight distribution, a learnable codebook optimization mechanism employing noise substitution...

Bao Tan Duy Huynh, Takashi Tsunakawa, Masafumi Nishida · 0 citations
#machine learning Preprint Sep 2026

Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures

Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. We evaluate PTQ across TTS architectures under one protocol with three core models, weight and activation ablations of eight more, and two held-out models quantized blind. Four-bit...

Se Un Park, Yutae Kim, Junyoung Park · 0 citations
Sep 2026

Adaptive multi-bit progressive quantization for stable training of binary neural networks

This novel progressive quantization framework combines multi-bit assistive teacher models with self-knowledge distillation to stabilize BNN training and integrates matching structured pruning with an asymmetric Binary Weight Network scaling factor, thereby reducing quantization errors while maintaining hardware efficie...

Jie Xu, Wonjun Hwang, Hyunsouk Cho · 0 citations
Preprint Aug 2026

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

Across five models from 1B to 32B, SoftWater outperforms the released WaterSIC quantizer at matched head rates on 59 of 60 test points, using none of that pipeline's refinements and cutting head-induced KL by $6.5\times$--$8.3\times at 2 bits.

Joao V. Cavalcanti, Ashia C. Wilson · 2 citations
#artificial intelligence Preprint Sep 2026

All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs

All for 1-Bit (AF1) is proposed, a genuine 1-bit PTQ framework for LLMs that consistently outperforms existing binarization-based PTQ methods in perplexity and zero-shot accuracy, providing a practical path toward deployable genuine 1-bit compression for LLMs.

Zhi-Xiong Zhao, Zu-Kang Xu, Guang-Yu Sun et al. · 1 citation
#machine learning Preprint Sep 2026

Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution

Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise,...

Ning-Feng Yang, T. Aamodt · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.