Skip to content
Open access

A pronunciation-sensitive analysis of XTTS adaptation for modern standard Arabic synthetic diacritization length rescue and training dynamics

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 40 references

TL;DR

The findings support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.

Abstract

Modern Standard Arabic (MSA) text-to-speech evaluation is complicated by the omission of short vowels and case endings in ordinary writing: a system may remain lexically intelligible while realizing a linguistically inappropriate pronunciation. This study examines XTTS v2 adaptation through matched raw, fully synthetic-diacritized, partially diacritized, and punctuation-guided length-rescue conditions. The retained Arabic speech pool contains 99,140 audio segments (296.94 h, 24 kHz mono), while the controlled adaptations use fixed subsets of 0.805–4.245 h. Evaluation combines a 60-sentence diagnostic benchmark, canonical ASR-proxy WER/CER, 99 targeted case-ending judgments, blinded expert naturalness and pronunciation ratings, pairwise preferences, and checkpoint analysis at 1, 3, and 13 total epochs. In the canonical evaluation, the off-the-shelf model was the strongest benchmark-wide reference (WER 0.224; CER 0.071), and every controlled adapted condition had a higher modeled overall error rate after Holm correction. Rescue-M produced the lowest descriptive case-sensitive CER (0.065), narrowly below B0 (0.067), but none of its planned case-sensitive contrasts remained statistically distinguishable after correction. An unseen explicit token substantially increased WER and CER (rate ratios 2.159 and 3.515). Validation loss declined across the observed checkpoints, whereas external error trajectories were non-monotonic and configuration-dependent. The independent evaluator produced a different case-ending ordering from R1, while pronunciation ratings favored B0 more consistently than naturalness or case-ending scores. The findings therefore support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.

Read PDF

Similar papers

#natural language process... Preprint Sep 2026

CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech span...

Bashar Talafha, Samar M. Magdy, Aisha Alansari et al. · 0 citations
Open access Sep 2026

Evaluation of Azure Pronunciation Assessment Tool for Arabic-Speaking Learners

Accurate reading aloud is crucial for Arabic-speaking learners because minor errors can alter meaning or hinder understanding. This study evaluated the Microsoft Azure Pronunciation Assessment (PA) tool for modern standard Arabic using the KSU Arabic Speech Database. We focused on male speakers recorded in silent rooms...

J. Zainaddin, Mohammad Amro · 0 citations
Preprint Aug 2026

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system differences as recognition errors. Jointly normalizing reference and hypothesis avoids this, but...

H. Asadpour · 0 citations
Open access Sep 2026

Jawhar: Optimized Morphological Analysis and Contextual Reranking for Arabic Part-of-Speech Tagging

Part-of-speech (POS) tagging in Arabic is hard because its rich root-and-pattern morphology and the absence of short vowels make one unvoweled string compatible with many categories. This paper presents Jawhar, a hybrid framework that couples a high-performance morphological analyser with contextual reranking using a p...

Mohamed Bouzahir, A. A. Abdelouahad, M. Nabil · 0 citations
#natural language process... Preprint Oct 2026

SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic

Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotat...

Ben Sapirstein, Roy Mattar, Guy Mor-Lan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.