Preprint
Aug 2026
Iterative Self-Learning for Expressive Text-to-Speech Synthesis
In the most data-scarce conditions, ISL-trained models outperform single-pass pseudo-labeling and further approach fully supervised performance, demonstrating that gradient-based ISL is an effective solution to expressive label scarcity in low-resource TTS.
Nicholas Sanders, G. Henter, Simon King et al.
· 0 citations