Back to feed
Conference

Data Saturation in Low-Resource TTS Fine-Tuning

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 24 references

Abstract

When fine-tuning large language models, the assumption is that more data means better results. This principle is often extended to text-to-speech (TTS) fine-tuning, yet remains underexplored, particularly for low-resource languages where high-quality data is difficult to obtain. In this work, the "more data is better" phenomenon is investigated for Turkish TTS using XTTS v2 by incrementally increasing training data. Speech quality is evaluated using multiple metrics, including UTMOS, NISQA, and an LLM-based TTS evaluation framework. For the LLM-based evaluation, a multimodal language model (Gemini) was prompted to assess Turkish-specific pronunciation, naturalness, and synthesis artifacts on a calibrated 1-10 scale. Results suggest that for TTS fine-tuning on non-mainstream languages, modest data investments may achieve near-optimal quality.

View source