Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on uncropped video suffer from weak temporal a...
Yu-Sheng Dai, Kangdi Wang, Baolong Gao et al.· 1 citation
Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving...
DuoTok is presented, a source-aware dual-track music tokenizer for vocal-accompaniment generation based on staged disentanglement, suggesting that tokenizer design is a core modeling problem for multi-track music generation, beyond compression alone.
Rui Lin, Zhiyue Wu, Jia-He Lei et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.