Skip to content

Author

Yusheng Dai

We have 4 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects

Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on uncropped video suffer from weak temporal a...

Yu-Sheng Dai, Kangdi Wang, Baolong Gao et al. · 1 citation
Preprint Aug 2026

Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction

Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving...

Kangdi Wang, Yu-Sheng Dai, Jin Xu · 0 citations
Preprint Sep 2026

Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visua...

Weiqiang Wang, Zhuo-Kun Chen, Yu-Sheng Dai et al. · 0 citations
Preprint Aug 2026

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

This work introduces TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification and significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as c...

Zehua Chen, Junyou Wang, Yuxuan Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.