This work evaluates speech and audio LLMs as a controlled call-decision problem, and finds that agreement and stacking features improve weaker selectors but do not beat the strongest no-call control.
Abstract
Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence and the need to call a generative audio model. We evaluate this distinction as a controlled call-decision problem. For each example, a policy chooses among keeping a transcript label, using encoder evidence from Contrastive Language-Audio Pretraining (CLAP), Audio Spectrogram Transformer (AST), or WavLM, and calling Qwen2-Audio, Qwen2.5-Omni, or MOSS-Audio; the decisive ablation removes all generative actions while keeping the selector and development protocol fixed. On VocalSound, the transcript-only representation reaches 0.296 accuracy, so it is insufficient for this label set. Yet supervised CLAP and WavLM controls reach 0.850 and 0.854 with no generative audio calls. A selector with generative actions reaches 0.925 accuracy using 12.5% calls, compared with 0.921 for the matched no-call selector (paired difference 0.004; 95% CI [-0.025, 0.033], including zero). Agreement and stacking features improve weaker selectors but do not beat the strongest no-call control. For known-task endpoint-value statements, the relevant quantity is the marginal value of the generative call after transcript and encoder evidence have already been used.
The Speech-Unsupported Rejection Evaluation Challenge (SURE-Challenge) is defined for this admission step and identifies a pre-generation error mode missed by answer-only scoring.
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whisper-Tiny and Whisper-Small with EnCodec...
Song-ha Jo, Sehyun Lee, Soyoon Kim et al.· 0 citations
Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary.
Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: producing feedback that identifies problems and suggests corrective actions grounded in the input. We introduce VocalCoachBench, a benchmark for e...
Hayeon Bang, Hounsu Kim, Wonil Kim et al.· 0 citations
Tffic-Audio is presented, a general speech deepfake detection system designed for comprehensive evaluation environment that achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard.
Wan Lin, Li Wang, Jin-Dong Wang et al.· arXiv.org· 1 citation
This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization, and finds that competitive SASV performance can be achieved through several distinct routes.
Sofya Savelyeva, Mariia Perunova, E. Kushnir et al.· arXiv.org· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.