Skip to content

Author

Peng Zhang

We have 4 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Video = World + Event Stream

We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.

Lianghua Huang, Zhigang Wu, Yupeng Shi et al. · 2 citations
Aug 2026

Towards Heterogeneous-Degradation-Robust Image Fusion with Controllable Generative Modulation.

Enhancing degradation robustness is essential for deploying image fusion techniques in real-world dynamic scenes. However, most existing methods either handle a single degradation type or assume fixed multi-degradation settings, making them insufficient for dynamically heterogeneous and composition ally complex degradations in practice. Moreover, they often fail to recover the semantics of salient scene targets when these targets are degraded or missing, leading to weakened semantic representation and reduced target saliency. To address these challenges, we propose DuS-DiFuse, a robust dual-stream latent diffusion framework composed of a diffusion fusion unit and a generative modulation unit. In the diffusion fusion unit, we fine-tune a CLIP visual encoder on multi-source data to perceive degradation types and severities, and employ latent diffusion to uniformly model multi-type, cross-level degradations with varying parameters. A Groupwise Fusion Control Module (GFCM) is further embedded into the latent degradation-removal process, enabling joint modeling of dynamic degradation removal and multimodal information fusion. In the generative modulation unit, pretrained latent diffusion priors are used to remodulate the initial fusion results, enabling controllable semantic restoration and generative enhancement, thereby improving target saliency and overall visual quality. To preserve fine-grained details during latent-to image reconstruction, we introduce a Detail-Restoration Fidelity Module (DRFM), which constrains texture reconstruction by jointly leveraging multi-level skip features from multiple source images and enhances structural fidelity in the fused results. Extensive experiments on multiple fusion datasets demonstrate that DuS-DiFuse achieves leading fusion performance, exhibits strong robustness to heterogeneous degradations, generalizes well across fusion tasks, and supports effective controllable generative modulation.

Lei Cao, Hao Zhang, Peng Zhang et al. · 0 citations
Preprint Aug 2026

Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

TetherMem is introduced, a training-free, query-aware spatiotemporal memory router for frozen video generators that separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject history and stale backgrounds.

Chen Li, Peng Zhang, Han-Yu Zhou et al. · 0 citations
Jul 2026

Wan-Streamer v0.2: Higher Resolution, Same Latency

Wan-Streamer v0.2 is presented, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model, which raises the interactive output stream from 192x336 to 640x368 while preserving approximately 200 ms model-side signal-to-signal latency at 25 FPS.

Lianghua Huang, Zhigang Wu, Yupeng Shi et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.