Skip to content
Preprint

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

A reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis, which reveals substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge.

Abstract

Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However, comprehensive evaluation methodologies that jointly assess perceptual quality, speaker similarity, and acoustic fidelity across diverse speech domains remain limited, particularly for low-resource and underrepresented languages. This paper presents a reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis. The proposed framework integrates complementary subjective and objective evaluation protocols and is demonstrated through a comprehensive case study on a representative low-resource language spanning four speech domains: Formal, Conversational, Literary/Storytelling, and Emotional. Four state-of-the-art TTS systems -- Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, and Google Gemini TTS -- are evaluated using MUSHRA listening tests, ABX discrimination tests, speaker similarity scoring with Resemblyzer, and acoustic analyses based on mel-cepstral distortion (MCD) and F0 RMSE over 960 audio pairs. Results reveal substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge (mean MCD 12.03 dB; mean F0 RMSE 889 cents), while conversational speech achieves the highest overall acoustic fidelity. Beyond the empirical findings, this work provides a reproducible evaluation framework, publicly releasing evaluation scripts, result tables, and executable Colab notebooks to support standardized benchmarking and future research on TTS evaluation for low-resource languages.

View source

Similar papers

Conference Jul 2026

Speech-to-Speech Translation Evaluation: A Systematic Analysis and Multi-Dimensional Metrics

Evaluating speech-to-speech translation (S2ST) systems remains challenging due to the multi-dimensional nature of speech, encompassing semantic accuracy, acoustic quality, and speaker characteristics. Existing approaches rely on either textbased or speech-based metrics, each capturing only partial aspects of translation quality. In this work, we conduct a systematic empirical analysis of S2ST evaluation metrics across multiple language pairs (fr-en, es-en, de-en, hi-en) using SeamlessM4T-v2 translations generated from the FLEURS dataset. We analyze ngram, neural text, and speech-embedding metrics at both corpus and sentence levels. Our analysis shows that text-based metrics are sensitive to linguistic variation but depend on automatic speech recognition (ASR), while speech-based metrics provide more consistent scores yet exhibit limited discriminative ability. Based on these observations, we propose a multi-dimensional evaluation framework that jointly assesses semantic adequacy and acoustic naturalness. The proposed framework improves semantic alignment and achieves strong correlation with speech naturalness across language pairs, enabling more reliable evaluation of S2ST systems.

Lalaram Arya, Mrinmoy Bhattacharjee, S. R. Mahadeva Prasanna · 0 citations
Open access Aug 2026

Speech-to-text model comparison using XLS-R, XLSR-53, and Wav2Vec 2.0

Findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian and provide practical guidance for deploying ASR systems in low-resource language scenarios.

J. Hebert, Amalia Zahra · 0 citations
Preprint Aug 2026

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct"naturalness"into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

Oluwanifemi Bamgbose, Simon Rosen, J. Shah et al. · 0 citations
Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Conference Jul 2026

QMOS: Qwen-Based MOS Prediction for TTS

Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Qwen2-Audio large audio language model and combines them with layer-weighted WavLM SSL representations through a learned cross-attention fusion. This design captures both high-level semantic naturalness and lowlevel acoustic distortions in a unified model. On two standard benchmarks, SOMOS and BVCC, QMOS achieves competitive system-level SRCC of 0.923 and 0.921 on SOMOS and BVCC, respectively, while using no system-ID conditioning or listener embeddings. Cross-domain evaluation yields a competitive SRCC of 0.702, suggesting the learned representations generalise across acoustic domains.

Ravi Sastry Kolluru, S. Devarakonda, S. Radhe Shyam Salopanthula et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.