Skip to content

Author

Xinggang Wang

We have 8 of 315 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Rethinking Representations for World-Action Modeling

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e...

Hao-Yi Jiang, Liu Liu, Xin-Jiang Wang et al. · 0 citations
Preprint Sep 2026

ReDrive: Shaping Representations with World Modeling for End-to-End Driving

Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by...

Yue-Ting Zhu, Shao-Yu Chen, Yue-Hao Song et al. · 0 citations
Preprint Sep 2026

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The...

Hong-Yuan Tao, Xing-Gang Wang, Liang-Hui Zhu et al. · 0 citations
Preprint Aug 2026

EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass

Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and achieves competitive open-vocabulary performance through exhaustive vocabulary traversal, yet suffers from prohibitive computational overhead as target categories scale. I...

Hao Peng, Yong-Kang Li, Zhao-Xiang Liu et al. · 0 citations
Preprint Sep 2026

ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We...

Si-Nuo Wang, Zichong Gu, Yu-Han Huang et al. · 0 citations
Preprint Aug 2026

DreamWAM: Beyond RGB Future Prediction for World Action Models

DreamWAM is introduced, which reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics, showing that robust world-action learning depends not only on predicting the future, but on representing it in a fo...

Shanglin Yuan, Weiheng Zhao, Xin Shi et al. · 4 citations
Preprint Aug 2026

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Faster-WAM introduces a sparse future-conditioning framework that computes future representations once and selectively reuses them throughout action denoising, and proposes SparseMoT to replace ubiquitous layer-wise fusion with selective video-action interaction at a compact subset of network stages, and Interval KV-Fu...

Weiheng Zhao, Haoyi Jiang, Xin Shi et al. · 10 citations
Preprint Aug 2026

Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

This work reformulates the video diffusion sampling as a frame-indexed stochastic process over noise levels, and constructs a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling.

Yue-Ting Zhu, Yue-Hao Song, Kai-Chen Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.