Skip to content

Author

Junhao Zhuang

We have 6 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint op...

Zi-Han Su, Junhao Zhuang, Yao-Wei Li et al. · 0 citations
Preprint Sep 2026

Where and When to Force: Routed Forcing for Streaming Avatars

Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes...

Zi-Han Su, Si-Wen Lu, Junhao Zhuang et al. · 0 citations
Preprint Aug 2026

EchoWM: Open and Enterable Omnimodal World Models

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...

Song-Chun Zhang, Yao-Wei Li, Junhao Zhuang et al. · 5 citations · ⚡1
Jul 2026

Self Gradient Forcing: Native Long Video Extrapolation

Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory.

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian et al. · 5 citations
Preprint Aug 2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Nan Duan, Hao-Yang Huang, Wei-Yang Jin et al. · 3 citations · ⚡1
Jul 2026

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring

An efficient generative framework designed to improve in-the-wild robustness under diverse real capture conditions and demonstrate strong perceptual quality, semantic fidelity, and temporal consistency on unseen videos, as well as improved robustness in downstream 3D reconstruction under severe motion blur.

Renbiao Jin, Ming-Hsuan Yang, Yutian Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.