Skip to content

Author

Lvmin Zhang

We have 10 of 27 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

EchoWM: Open and Enterable Omnimodal World Models

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...

Song-Chun Zhang, Yao-Wei Li, Junhao Zhuang et al. · 5 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Training Object Permanence in World Models

This work introduces WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories, and builds Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's c...

Hao-Tian Zhang, Feng-Yuan Yu, Dezhi Luo et al. · 0 citations
Preprint Sep 2026

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

VC-Attention is proposed, a training-free low-bit attention framework that addresses diffusion Transformers and quantization scale by pairing Value smoothing with a fused probability Cast, and improves fidelity over low-bit baselines.

Xing-Yang Li, Dong-Yun Zou, Shi-Ning Zhang et al. · 1 citation · ⚡1

Mode Seeking meets Mean Seeking for Fast Long Video Generation

This paper proposes a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence based on a unified representation via a Decoupled Diffusion Transformer, and closes the fidelity-horizon gap by jointly improving local sharpness, motion and long-range consistency.

Shengqu Cai, Weili Nie, Chao Liu et al. · 12 citations · ⚡2
#machine learning Preprint Sep 2026

RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives

We present'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materia...

Chong Zeng, Yue Dong, P. Peers et al. · 0 citations
Preprint Aug 2026

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory. This deficiency can be studied along two critical aspects: object permanence, the ability to precisely reproduce the appearance of objects upon re-entry; and memory capacity, the ability to process ult...

Bowen Xue, Brandon Y. Feng, Chenguo Lin et al. · 0 citations
#artificial intelligence Review Mar 2026

View-oriented Conversation Compiler for Agent Trace Analysis

View oriented Conversation Compiler is proposed, namely View oriented Conversation Compiler, which lexes, parses, and lowers a raw JSONL log into three views based on one intermediate representation, which demonstrates the effectiveness of trace format as an important component of context engineering infrastructure.

Lvmin Zhang, Maneesh Agrawala · 1 citation
#machine learning Preprint Aug 2026

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.

Junhua Xu, Rui-Si Wang, Fanyi Pu et al. · 5 citations
Jul 2026

Masked Visual Actions for Unified World Modeling

Masked Visual Actions is introduced, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video that supports inverse modeling by synthesizing robot motion from desired object motion.

Hadi Alzayer, Wenlong Huang, Haonan Chen et al. · 3 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.