This work introduces Self-Contrastive Residual Alignment (SCRA), a one-pass inference-time edit that learns a low-rank linear corrector for the prompt-induced residual shift and slightly but significantly improves over the audio-only baseline on difficult samples.
Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary.
A mechanistic analysis of paralinguistic information in four open source models using the Expresso dataset with controlled speaking styles identifies a gap between what models encode and what they use, highlighting a key limitation in current audio language models.
Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or cont...
Yihui Fu, Zhengyang Li, Tim Fingscheidt· 0 citations
As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LA...
Fernando López, A. Ayala, Guillermo Segovia et al.· arXiv.org· 0 citations
In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inf...
Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al.· 0 citations
Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so...