Dec 2025
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
StellarTTS is introduced, a novel mobile-optimized NAR TTS framework based on a sparse temporal embedding strategy, enabling granular control of phoneme duration, pronunciation, and prosody and a semantic-aware codec that facilitates efficient single-stage decoding.
Kaicheng Luo, Xuefei Gong, Yu-Tao Sun et al.
· Automatic Speech Recognition... · 0 citations