2026· Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada· 0 citations
TL;DR
It is demonstrated that deep bidirectional interaction between quantized acoustic streams and semantic contexts is essential for mitigating ASR error propagation and achieving robust Chinese SNER.
Abstract
Conventional Speech Named Entity Recognition (SNER) typically relies on cascaded ASR (Automatic Speech Recognition)+NER (Named Entity Recognition) pipelines, which are hindered by error propagation and the underutilisation of acoustic cues. We propose an end-to-end Chinese SNER framework using Residual Vector Quantisation (RVQ) and deep acoustic--semantic fusion. The model extracts speech representations via a frozen Wav2Vec2-XLSR encoder, employing an RVQ-based bottleneck to reconstruct continuous quantized features that regularize the acoustic space and preserve semantic content.
A Transformer decoder, trained with a joint CTC-attention objective, performs transcription while a gated deep-fusion mechanism integrates an external GPT model for linguistic consistency. For NER, a bidirectional multimodal fusion module aligns acoustic and semantic features before a GlobalPointer head performs span-level prediction. Experiments on AISHELL-NER and CNERTA yield F1-scores of 90.91\% and 81.44\%, respectively, outperforming pipeline, multimodal, and E2E baselines. These results demonstrate that deep bidirectional interaction between quantized acoustic streams and semantic contexts is essential for mitigating ASR error propagation and achieving robust Chinese SNER.
An AS-Split Conformer–Mamba framework that decouples local and global modeling into two explicit stages and achieves consistent improvements over strong baselines is proposed.
Lulu Qin, Xuan Fu, Mingchen Sun et al.· Electronics· 0 citations
End-to-end on-device Automatic Speech Recognition (ASR) systems have demonstrated remarkable accuracy and efficiency in recent years. However, challenges persist in correctly transcribing infrequent named entities (e.g., geographical locations, business entities, person names, etc.) and handling diverse user accents, w...
Kiranmayi Gandikota, Anunay Katare, C. Pandey et al.· International Conference on...· 0 citations
Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Q...
Ravi Sastry Kolluru, S. Devarakonda, S. Radhe Shyam Salopanthula et al.· International Conference on...· 0 citations
ParaASR is introduced, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step and shows that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal...
Experiments on two 30-hour datasets covering the low-resource Sinitic varieties Taiwanese Hokkien and Hakka show that SAMA-ASR improves over acoustic, prior prompt-based, and semantic-only translation-guided baselines and remains effective in practical automatic semantic-anchor settings; translator-capacity analyses sh...
Kuan-Tang Huang, Cheng-Yeh Yang, Chien-Chun Wang et al.· 0 citations
Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training procedure guided by text predictability. Star...
Hanlin Zhang, Da-Xin Tan, De-Hua Tao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.