Skip to content

Author

Xinyuan Chen

We have 5 of 38 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers

Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry...

Yu-Tong Wang, Xing-Tong Ge, En-Huai Liu et al. · 1 citation
Preprint Oct 2026

VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation

Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can prov...

Yu-Tong Wang, Xing-Tong Ge, En-Huai Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies

Every valid trajectory can teach what physically happened, while the actor is deployed only under the condition associated with relatively better actions, in \method, a direct world-action policy that combines outcome-agnostic predictive learning with outcome-aware policy improvement.

Bo-Tong Zhao, Fangjie Yu, Tim Yu et al. · 0 citations
Open access Aug 2026

Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model

Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model, is proposed.

Haoning Yang, Xinyuan Chen, Yaohui Wang et al. · 1 citation
Preprint Jul 2026

DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking

DeforM is proposed, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions, and introduces a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks.

Yunyi Li, Yu Qiao, Yaohui Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.