Preprint
Jul 2026
Fr\'echet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Speech Representation Fr'echet Distance loss (SR-FD), a training-time distributional regularizer for tokenizer-free flow-matching autoregressive TTS, is proposed, an intelligibility-improving distributional regularizer for few-step TTS.
Ho-Lam Chung, Kuan-Po Huang, Bo-Ru Lu et al.
· 0 citations