GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction, and introduces a fixed-receptive-field convolutional encoder that reduces the respective prediction errors.
Zi-Tao Liang, Chang Gao
· 0 citations