Skip to content

HPQ: A Hybrid Framework for Joint Pruning and Quantization of Self-Supervised Speech Models

2026 · IEEE Signal Processing Letters · Vol 33, pp. 3058-3062 · 0 citations · 39 references

Abstract

Despite achieving state-of-the-art accuracy in speaker verification (SV), large-scale self-supervised learning (SSL) speech models remain difficult to deploy on edge devices because of their computational and memory demands. Existing compression approaches improve efficiency through pruning and quantization, but usually optimize them sequentially and thus overlook their interaction. In this paper, we propose HPQ, a framework that jointly optimizes differentiable structured pruning and learnable quantization in a single fine-tuning stage. By integrating <inline-formula><tex-math notation="LaTeX">$L_{0}$</tex-math></inline-formula> regularization with Learned Step Size Quantization (LSQ) into a unified objective, HPQ enables the network to co-adapt its architecture to quantization noise. Experiments on VoxCeleb demonstrate that HPQ establishes a new Pareto frontier: an 8-bit, 70% sparse WavLM model achieves a <inline-formula><tex-math notation="LaTeX">$13\times$</tex-math></inline-formula> reduction in model size and a <inline-formula><tex-math notation="LaTeX">$15\times$</tex-math></inline-formula> reduction in bit-operations, with only a 0.23% absolute EER degradation compared with the full-precision baseline. We further show that larger backbones and more diverse training data improve robustness under aggressive compression across both WavLM and W2V-BERT.

View source

Similar papers

OJBKQ : Objective-Joint Babai–Klein-Based Quantization ∗

OJBKQ is proposed, a layer-wise PTQ method that formulates weight quantization as a joint optimization problem over activations and weights, yielding a multiple-right-hand-side box-constrained integer least squares (BILS) problem per layer.

Xinyu Wang, Ziyu Zhao, Peng Lu et al. · 0 citations
Jul 2026

CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

Ultra-lightweight models are essential for the deployment of deep learning-based speech enhancement algorithms on edge devices. Although recent approaches have achieved a certain balance between computational complexity and performance, pushing the complexity limits further demands more sophisticated designs. In this letter, we propose CoFi-Lite, a highly efficient model that decouples spectral modeling into coarse- and fine-grained streams. By leveraging two parallel and symmetric encoder-decoder paths, it simultaneously extracts full-band envelopes and low-frequency details for complementary enhancement. In addition, a novel Cross-Path Fusion (CPF) module is introduced to bridge the distinct paths, facilitating efficient feature interaction. Remarkably, CoFi-Lite requires extremely low computational resources, featuring only 12.87 M MACs/s and 83.12 k parameters. Experimental results demonstrate that our proposed model outperforms the ultra-lightweight baseline GTCRN while requiring only 40.26% of its computational complexity. Its scaled-up variant also delivers performance on par with that of the SOTA ultra-lightweight model AdaptCRN alongside a 19.34% reduction in computational cost.

Leyan Yang, Dahan Wang, Xiaobin Rong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied after every column block (128 columns). By coupling this fine-grained correction with a sliding window mechanism for smooth cross-layer transitions, REAL-Q effectively mitigates error propagation across the network. On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, REAL-Q reduces end-to-end KL divergence by up to ~49% relative to state-of-the-art globally-guided methods.

Qian Zhang, Yao-Ming Li, Zheng Tan et al. · 0 citations
Jul 2026

C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

C-PTQ is proposed, a unified channel-wise PTQ method that harmonizes task-specific loss perturbation and quantization error and achieves state-of-the-art performance without auxiliary modules like LoRA, thereby maintaining high efficiency.

Jiameng Li, Han Zhou, M. Blaschko · 0 citations
Preprint Aug 2026

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.

A. Shukla, R. Thakur, Aryan Das et al. · 0 citations
Preprint Aug 2026

SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features

SQuaT (Student-Aware Quantized Teacher Features), a label-free QAT framework with KD that theoretically eliminates this lower bound on the distillation loss by applying the student's quantization parameters to quantize the teacher's features during distillation is proposed.

H. Lee, Hyeonsik Jo, Jinwook Chung et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.