Skip to content

Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluatio

Aug 2026 · 0 citations · 31 references
Computer Science Engineering

TL;DR

This work evaluates speech and audio LLMs as a controlled call-decision problem, and finds that agreement and stacking features improve weaker selectors but do not beat the strongest no-call control.

Abstract

Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence and the need to call a generative audio model. We evaluate this distinction as a controlled call-decision problem. For each example, a policy chooses among keeping a transcript label, using encoder evidence from Contrastive Language-Audio Pretraining (CLAP), Audio Spectrogram Transformer (AST), or WavLM, and calling Qwen2-Audio, Qwen2.5-Omni, or MOSS-Audio; the decisive ablation removes all generative actions while keeping the selector and development protocol fixed. On VocalSound, the transcript-only representation reaches 0.296 accuracy, so it is insufficient for this label set. Yet supervised CLAP and WavLM controls reach 0.850 and 0.854 with no generative audio calls. A selector with generative actions reaches 0.925 accuracy using 12.5% calls, compared with 0.921 for the matched no-call selector (paired difference 0.004; 95% CI [-0.025, 0.033], including zero). Agreement and stacking features improve weaker selectors but do not beat the strongest no-call control. For known-task endpoint-value statements, the relevant quantity is the marginal value of the generative call after transcript and encoder evidence have already been used.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs

Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whisper-Tiny and Whisper-Small with EnCodec...

Song-ha Jo, Sehyun Lee, Soyoon Kim et al. · 0 citations
Jul 2026

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary.

Jiachen Qian, Junyu Li · 0 citations
Preprint Aug 2026

VocalCoachBench: Benchmarking Audio-Language Models on Expert Feedback for Singing

Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: producing feedback that identifies problems and suggests corrective actions grounded in the input. We introduce VocalCoachBench, a benchmark for e...

Hayeon Bang, Hounsu Kim, Wonil Kim et al. · 0 citations
Jul 2026

Teffic-Audio: Tell Fact from Fiction

Tffic-Audio is presented, a general speech deepfake detection system designed for comprehensive evaluation environment that achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard.

Wan Lin, Li Wang, Jin-Dong Wang et al. · 1 citation
Jul 2026

Large Audio Language Models for Spoofing-Aware Speaker Verification

This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization, and finds that competitive SASV performance can be achieved through several distinct routes.

Sofya Savelyeva, Mariia Perunova, E. Kushnir et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.