Skip to content
Open access

Audio-guided articulatory distillation for multilingual visual speech recognition with large language model decoding

Aug 2026 · Multimedia tools and applications · Vol 85 · 0 citations · 37 references

TL;DR

An audio-guided distillation framework for multilingual VSR that exploits synchronized audio-visual speech during training while preserving visual-only inference is introduced, and results further confirm the contribution of LLM decoding, teacher initialization, and audio-guided distillation to visual-only recognition performance.

Abstract

Visual Speech Recognition (VSR) remains challenging in multilingual and low-resource settings due to visual ambiguity, limited annotated data, and weak cross-lingual generalization. This paper introduces an audio-guided distillation framework for multilingual VSR that exploits synchronized audio-visual speech during training while preserving visual-only inference. The proposed architecture consists of an audio-visual teacher that learns articulation-aware representations from aligned acoustic and video streams, and a visual-only student trained to approximate the teacher through representation- and decoder-level distillation. To improve transcription under ambiguous visual evidence, continuous speech representations are projected into the embedding space of a pretrained large language model through a lightweight adaptation module, enabling language-conditioned decoding without full LLM retraining. We further introduce RoVSR-II, an extended Romanian in-the-wild VSR corpus comprising approximately 250 h of audiovisual speech, designed to support evaluation in an underrepresented language. Experiments on mTEDx demonstrate consistent improvements over existing multilingual VSR methods across Latin-script languages in terms of Word Error Rate (WER) and Character Error Rate (CER). Additional evaluation on RoVSR-II shows that the proposed model supports zero-shot transfer to Romanian and substantially reduces both WER and CER after parameter-efficient supervised adaptation. Ablation results further confirm the contribution of LLM decoding, teacher initialization, and audio-guided distillation to visual-only recognition performance. © 2017 Elsevier Inc. All rights reserved.

Read PDF

Similar papers

Preprint Aug 2026

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance.

Yexing Du, Kaiyuan Liu, Youcheng Pan et al. · 0 citations
Preprint Aug 2026

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

FireRedAudio is introduced, a general-purpose audio language model with a shared 9B-parameter LLM that achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantial improvements over Ming-UniAudio-Edit in both semantic and acoustic speech editing.

Junjie Li, Xuelong Geng, Kun Xie et al. · 2 citations · ⚡1
Jul 2026

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

ParaASR is introduced, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step and shows that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.

Qing-Jian Lin, Yuxin Li, Haoyang Zhang et al. · 1 citation
Preprint Aug 2026

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and their temporal relationships. Existing omni models typically introduce dedicated audio encoders and rely on expensive audio-video-text training, tightly coupling omni capability to specific VLM backbones and potentially weakening their existing visual and reasoning abilities. This raises three questions: whether native omni training is necessary for every new VLM, whether speech-centric omni capability can be added while preserving the original backbone, and where richer acoustic representations remain essential. We introduce Training-Free Omni (TFO), a plug-and-play framework that converts a frozen VLM into a speech-centric omni model without architectural modification, or multimodal re-alignment. TFO uses Whisper to extract confidence-filtered, timestamped transcripts and routes them through the VLM's existing language interface, while leaving its visual pathway unchanged. Across matched comparisons with native omni models on 56 benchmarks and 21 languages, TFO is competitive on audio-visual understanding, improves average audio-only performance across all five model settings, and achieves substantial multilingual speech gains. Freezing the VLM also generally preserves stronger image/video understanding, visual grounding, coding, mathematical reasoning, and medical question answering than the corresponding native omni checkpoints. These results show that strong speech-centric omni understanding can often be obtained through modular audio-to-language routing rather than costly backbone-specific training.

Unknown authors · 0 citations
Preprint Sep 2026

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field toward realistic dialogue, we introduce Candor-LR, a conversational benchmark derived from the CANDOR corpus of 1,656 natural dyadic videoconferences. Our custom data preparation pipeline yields 713.5, 10.1, and 60.1 hours of training, validation, and test data, respectively. Evaluating pretrained AVSR models on Candor-LR reveals that audio-only accuracy drops sharply compared to LRS3, but visual cues compensate effectively, driving much larger performance gains on Candor-LR than on LRS3. Furthermore, training on this corpus significantly improves cross-domain robustness under both clean and noisy conditions, as its realistic conversational data captures broader audio-video features. We open-source our pipeline to ensure reproducibility, establishing Candor-LR as a challenging benchmark for conversational AVSR.

Unknown authors · 0 citations
Preprint Aug 2026

Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR

Large language model (LLM)-based automatic speech recognition (ASR) systems achieve strong performance on conventional speech data by leveraging powerful linguistic priors and multilingual capabilities. However, under challenging conditions, these priors can override acoustic evidence, resulting in unintended translation, instruction execution, repetition, or catastrophic deletion. We propose Likelihood-Constrained Acoustic Reranking (LCAR), a training-free decoding method that improves acoustic grounding while preserving support from the base model. At each decoding step, LCAR first retains tokens whose base-model likelihood falls within a margin of the greedy token, then reranks them using an acoustic compatibility score computed from attention-pooled audio embeddings and the existing LM head. By restricting acoustic intervention to plausible, model-supported alternatives, LCAR requires no additional training, external detector, reference transcript, or auxiliary model at inference. We evaluate LCAR on four LLM-based ASR systems using human-audited TTS and open-source speech challenge suites. At $\delta=0.60$, LCAR removes 38.8--57.1\% of detector-identified hallucination failures while largely maintaining WER/CER on standard open-source test sets.

Ji-Shen Kuang, Lin-Ru Zheng, Hongjin Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.