Skip to content
Preprint

Is Semantics Enough for Speech Mean Opinion Score Prediction?

Tian-Yu Lan Yu-Fei Shi Yang Ai Hong-Hao Sun Hui-Peng Du Zhen-Hua Ling
Sep 2026 · 0 citations · 27 references
Computer Science

Abstract

Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.