A unified theoretical framework is introduced that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow and motivates MGFlow, which models feature distributions with Gaussian mixtures at an adjustable granularity between...
Chi Zhang, Hao-Yan Shi, Yue-Yi Liu et al.· 2 citations
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework mi...
Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al.· 1 citation
Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is...
Chi Zhang, Hao-Yan Shi, Yueyi Liu et al.· 2 citations
This work introduces MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings and establishes MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.