CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Abstract
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners'speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak&Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.
As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LALMs across heterogeneous acoustic conditions and reasoning abilities. ESCUCHA comprises 1,000 human-curated questions paired with audio, totaling 162.9 hours sourced directly ``from the wild''rather than drawn from existing datasets, with durations ranging from a few seconds to over 80 minutes. The benchmark emphasizes reasoning, spanning 9 perceptual and 10 reasoning categories, and it captures linguistic diversity through multiple Spanish accents and non-normative speech. ESCUCHA further includes multi-audio questions, spoken questions, and audio instructions, and it flags which questions support open-ended evaluation. Benchmarking several state-of-the-art multimodal and speech models reveals substantial performance gaps relative to trained humans.
Fernando López, A. Ayala, Guillermo Segovia et al.· arXiv.org· 0 citations
Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Qwen2-Audio large audio language model and combines them with layer-weighted WavLM SSL representations through a learned cross-attention fusion. This design captures both high-level semantic naturalness and lowlevel acoustic distortions in a unified model. On two standard benchmarks, SOMOS and BVCC, QMOS achieves competitive system-level SRCC of 0.923 and 0.921 on SOMOS and BVCC, respectively, while using no system-ID conditioning or listener embeddings. Cross-domain evaluation yields a competitive SRCC of 0.702, suggesting the learned representations generalise across acoustic domains.
Ravi Sastry Kolluru, S. Devarakonda, S. Radhe Shyam Salopanthula et al.· International Conference on...· 0 citations
Evaluating speech-to-speech translation (S2ST) systems remains challenging due to the multi-dimensional nature of speech, encompassing semantic accuracy, acoustic quality, and speaker characteristics. Existing approaches rely on either textbased or speech-based metrics, each capturing only partial aspects of translation quality. In this work, we conduct a systematic empirical analysis of S2ST evaluation metrics across multiple language pairs (fr-en, es-en, de-en, hi-en) using SeamlessM4T-v2 translations generated from the FLEURS dataset. We analyze ngram, neural text, and speech-embedding metrics at both corpus and sentence levels. Our analysis shows that text-based metrics are sensitive to linguistic variation but depend on automatic speech recognition (ASR), while speech-based metrics provide more consistent scores yet exhibit limited discriminative ability. Based on these observations, we propose a multi-dimensional evaluation framework that jointly assesses semantic adequacy and acoustic naturalness. The proposed framework improves semantic alignment and achieves strong correlation with speech naturalness across language pairs, enabling more reliable evaluation of S2ST systems.
Lalaram Arya, Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna· International Conference on...· 0 citations
Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.
Findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian and provide practical guidance for deploying ASR systems in low-resource language scenarios.
J. Hebert, Amalia Zahra· Bulletin of Electrical Engin...· 0 citations
A reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis, which reveals substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge.
Ali B. Jafar, Amal Sarmad, Shifa Yousaf et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.