Jul 2026
Wonder: Video World Model Done Better
This work introduces a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence, regardless of actual context length.
Jiacong Xu, Hanwen Jiang, Zhixin Shu et al.
· arXiv.org · 3 citations
· ⚡1