The findings support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.
Abstract
Modern Standard Arabic (MSA) text-to-speech evaluation is complicated by the omission of short vowels and case endings in ordinary writing: a system may remain lexically intelligible while realizing a linguistically inappropriate pronunciation. This study examines XTTS v2 adaptation through matched raw, fully synthetic-diacritized, partially diacritized, and punctuation-guided length-rescue conditions. The retained Arabic speech pool contains 99,140 audio segments (296.94 h, 24 kHz mono), while the controlled adaptations use fixed subsets of 0.805–4.245 h. Evaluation combines a 60-sentence diagnostic benchmark, canonical ASR-proxy WER/CER, 99 targeted case-ending judgments, blinded expert naturalness and pronunciation ratings, pairwise preferences, and checkpoint analysis at 1, 3, and 13 total epochs. In the canonical evaluation, the off-the-shelf model was the strongest benchmark-wide reference (WER 0.224; CER 0.071), and every controlled adapted condition had a higher modeled overall error rate after Holm correction. Rescue-M produced the lowest descriptive case-sensitive CER (0.065), narrowly below B0 (0.067), but none of its planned case-sensitive contrasts remained statistically distinguishable after correction. An unseen explicit token substantially increased WER and CER (rate ratios 2.159 and 3.515). Validation loss declined across the observed checkpoints, whereas external error trajectories were non-monotonic and configuration-dependent. The independent evaluator produced a different case-ending ordering from R1, while pronunciation ratings favored B0 more consistently than naturalness or case-ending scores. The findings therefore support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech span...
Bashar Talafha, Samar M. Magdy, Aisha Alansari et al.· 0 citations
Accurate reading aloud is crucial for Arabic-speaking learners because minor errors can alter meaning or hinder understanding. This study evaluated the Microsoft Azure Pronunciation Assessment (PA) tool for modern standard Arabic using the KSU Arabic Speech Database. We focused on male speakers recorded in silent rooms...
J. Zainaddin, Mohammad Amro· Journal of Undergraduate Res...· 0 citations
Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system differences as recognition errors. Jointly normalizing reference and hypothesis avoids this, but...
Part-of-speech (POS) tagging in Arabic is hard because its rich root-and-pattern morphology and the absence of short vowels make one unvoweled string compatible with many categories. This paper presents Jawhar, a hybrid framework that couples a high-performance morphological analyser with contextual reranking using a p...
Mohamed Bouzahir, A. A. Abdelouahad, M. Nabil· Information· 0 citations
Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotat...
Ben Sapirstein, Roy Mattar, Guy Mor-Lan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.