Skip to content
Preprint

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

Jul 2026 · 0 citations · 40 references
Engineering

TL;DR

Results confirm that while models perform significantly better on seen samples, frozen SSL embeddings combined with deep Transformer encoders offer the most stable and scalable solution for universal speech quality assessment.

Abstract

Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. This study presents a comprehensive benchmarking of three architectural frameworks: Frozen Self-Supervised Learning (SSL-FRZ), Fine-Tuned SSL (SSL-FT), and a Video Vision Transformer (ViViT). Evaluation is conducted in two phases: Part I utilizes a consolidated corpus of 130,000 samples across 19 diverse datasets, while Part II focuses on a purified 17-dataset English-only corpus. To assess robustness, a systematic Leave-One-Dataset-Out (LODO) protocol is employed to quantify the generalization gap between seen and unseen distributions. Finally, the top-performing model is benchmarked against 18 state-of-the-art (SOTA) metrics using the ARECHO framework. Results demonstrate that an English-only purified corpus consistently yields higher predictive precision across all architectures. While SSL-FT achieves the highest performance on seen validation data, the SSL-FRZ model provides superior robustness on unseen distributions, achieving a competitive Mean Squared Error (MSE) of 0.36 on the URGENT 2024 benchmark-closely matching domain-optimized SOTA metrics (MSE 0.30). Although the ViViT architecture remains below SSL-based models in total capacity, it delivers stable results in English-only trials. LODO results confirm that while models perform significantly better on seen samples, frozen SSL embeddings combined with deep Transformer encoders offer the most stable and scalable solution for universal speech quality assessment. To support further research, the top-performing English-only SSL-Transformer model and weights are made publicly available via Hugging Face.

View source

Similar papers

Conference Jul 2026

QMOS: Qwen-Based MOS Prediction for TTS

Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Q...

Ravi Sastry Kolluru, S. Devarakonda, S. Radhe Shyam Salopanthula et al. · 0 citations
Open access 2026

A System-Level Analysis of Cross-Domain Generalization Failure in Audio Deepfake Detection

Audio deepfake detectors often report high accuracy on individual benchmarks, yet their reliability under domain shift remains largely untested. This study presents a system-level analysis of cross-domain generalization failure, evaluating five hybrid architectures (CNN–LSTM, TCN, TCN–LSTM, Conformer, and TCM-Conformer...

Saadin Oyucu, Bilgehan Arslan, Şeref Sağıroğlu · 0 citations
Preprint Sep 2026

Is Semantics Enough for Speech Mean Opinion Score Prediction?

Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper,...

Tian-Yu Lan, Yu-Fei Shi, Yang Ai et al. · 0 citations
Conference Jul 2026

MRAN-UNet: Physics-Informed Harmonic Frequency Attention for Multilingual Speech Enhancement

Deep learning speech enhancement models are trained without grounding in acoustic physics, and evaluations remain confined almost exclusively to English. We address both gaps with MRAN-UNet, which embeds Harmonic Frequency Attention (HFA) - a parameter-free module derived from the source-filter model that aggregates sp...

Satya Prasad Gaddamedi, Debolina Pramanik, Puja Bharati et al. · 0 citations
Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Open access Aug 2026

Generalization Analysis of YAMNet-DNN Architectures in Deepfake Audio Classification

The rapid advancement of speech synthesis and voice conversion technologies has increased the risk of audio deepfake attacks, necessitating robust and generalizable detection systems. This study proposes a deepfake audio detection framework that leverages pretrained YAMNet embeddings as a feature extractor, combined wi...

Hakam Dzakwan Diash, Dwi Arman Prasetya, Alfan Rizaldy Pratama et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.