Skip to content

Author

Nan Duan

We have 14 of 36 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Oct 2026

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint op...

Zi-Han Su, Junhao Zhuang, Yao-Wei Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation...

Ya-Feng Chen, Bo-Yan Dong, Yan-Kun Huang et al. · 0 citations
Preprint Sep 2026

Where and When to Force: Routed Forcing for Streaming Avatars

Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes...

Zi-Han Su, Si-Wen Lu, Junhao Zhuang et al. · 0 citations
Preprint Sep 2026

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and...

Shenghe Zheng, Wen-Bo Li, Ji-Yao Zhang et al. · 1 citation
Preprint Aug 2026

EchoWM: Open and Enterable Omnimodal World Models

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...

Song-Chun Zhang, Yao-Wei Li, Junhao Zhuang et al. · 5 citations · ⚡1

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Across six real-world scenarios, human raters prefer JoyAI-VL-Interaction over the in-app video-call assistants of Doubao and Gemini by a wide margin, and the first open, vision-driven interaction model released together with its training recipe, data, and complete deployable system.

Ding-Yu Yao, Jun Zhou, Chenxu Yang et al. · 10 citations
Preprint Sep 2026

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

A large-scale synthetic data production pipeline built on Unreal Engine for generating action-conditioned, multi-view video, and describes the system architecture, the implementation decisions that emerged from production failures, and the limitations of using perceptual quality proxies for world-model data curation.

Hao-Yu Wang, Song-Chun Zhang, Hao-Ran Li et al. · 0 citations
Preprint Sep 2026

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

FactoSR is presented, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection, and suggests that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

Yi-Jun Yang, Shenghe Zheng, Wen-Bo Li et al. · 0 citations
Preprint Aug 2026

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.

Xiong-Hao Wu, Yi-Jun Yang, Shi-Long Zhou et al. · 3 citations
Preprint Aug 2026

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

The method combines chunk-wise autoregressive adaptation, Source-Anchored Distribution Matching Distillation, and Long-Horizon Autoregressive Distillation to reduce train--inference mismatch, preserve source fidelity during two-step generation, and mitigate accumulated temporal drift.

Yi-Cheng Xiao, Wenxun Dai, Xinran Qin et al. · 5 citations
Jul 2026

Self Gradient Forcing: Native Long Video Extrapolation

Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory.

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian et al. · 5 citations
Jul 2026

Perceptual Flow Matching for Few-Step Generative Modeling

Perceptual Flow Matching supervises flow matching in a perceptual feature space using pretrained perceptual models, which substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality.

Chuyang Zhao, Yifei Song, Hongfa Wang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.