Skip to content

Author

Jiangyu Han

We have 2 of 22 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Jul 2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.

Shuai Wang, Zihan Qian, Ke Zhang et al. · 1 citation
2026

HPQ: A Hybrid Framework for Joint Pruning and Quantization of Self-Supervised Speech Models

Despite achieving state-of-the-art accuracy in speaker verification (SV), large-scale self-supervised learning (SSL) speech models remain difficult to deploy on edge devices because of their computational and memory demands. Existing compression approaches improve efficiency through pruning and quantization, but usually optimize them sequentially and thus overlook their interaction. In this paper, we propose HPQ, a framework that jointly optimizes differentiable structured pruning and learnable quantization in a single fine-tuning stage. By integrating <inline-formula><tex-math notation="LaTeX">$L_{0}$</tex-math></inline-formula> regularization with Learned Step Size Quantization (LSQ) into a unified objective, HPQ enables the network to co-adapt its architecture to quantization noise. Experiments on VoxCeleb demonstrate that HPQ establishes a new Pareto frontier: an 8-bit, 70% sparse WavLM model achieves a <inline-formula><tex-math notation="LaTeX">$13\times$</tex-math></inline-formula> reduction in model size and a <inline-formula><tex-math notation="LaTeX">$15\times$</tex-math></inline-formula> reduction in bit-operations, with only a 0.23% absolute EER degradation compared with the full-precision baseline. We further show that larger backbones and more diverse training data improve robustness under aggressive compression across both WavLM and W2V-BERT.

Junyi Peng, Lin Zhang, Jiangyu Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.