Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS
HybridEmo is introduced, a post-training framework that initializes both multi-emotion TTS tasks with SFT and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward.
Yan Zhou, Yunqi Hong, Yang Feng
· 1 citation