Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools, is introduced, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks.
Zhen-Yu Liu, Zhangquan Chen, Ke-Yi Chen et al.· 0 citations
ViP-Rig is a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones into a frozen pretrained autoregressive generator.
Zihan Qin, Ming-Ze Sun, Yifan Mao et al.· arXiv.org· 1 citation
D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-po...
Junhao Chen, Mingjin Chen, Henghaofan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.