We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...
This work introduces WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories, and builds Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's c...
Hao-Tian Zhang, Feng-Yuan Yu, Dezhi Luo et al.· 0 citations
VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.
This paper proposes a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence based on a unified representation via a Decoupled Diffusion Transformer, and closes the fidelity-horizon gap by jointly improving local sharpness, motion and long-range consistency.
Shengqu Cai, Weili Nie, Chao Liu et al.· arXiv.org· 12 citations· ⚡2
We present'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materia...
Chong Zeng, Yue Dong, P. Peers et al.· 0 citations
Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory. This deficiency can be studied along two critical aspects: object permanence, the ability to precisely reproduce the appearance of objects upon re-entry; and memory capacity, the ability to process ult...
Bowen Xue, Brandon Y. Feng, Chenguo Lin et al.· 0 citations
View oriented Conversation Compiler is proposed, namely View oriented Conversation Compiler, which lexes, parses, and lowers a raw JSONL log into three views based on one intermediate representation, which demonstrates the effectiveness of trace format as an important component of context engineering infrastructure.
This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.
Junhua Xu, Rui-Si Wang, Fanyi Pu et al.· 5 citations
Masked Visual Actions is introduced, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video that supports inverse modeling by synthesizing robot motion from desired object motion.
Hadi Alzayer, Wenlong Huang, Haonan Chen et al.· arXiv.org· 3 citations· ⚡1
FlexComposer is proposed, a unified framework that standardizes video compositing as a trajectory-guided conditional generation task, enabling the seamless integration of both static images and dynamic footage.
Song-Chun Zhang, S. Guo, Xianghao Kong et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.