A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and recent 3D geometric foundation models. Our key observation is that no single estimator is optimal for diverse challenges, such as wide baselines, lack of texture, appearance...
Zhi-Ning Gu, Shang-Jie Du, Wei-Min Qiu et al.· 0 citations
Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-tho...
This training-free method improves the performance of both U-Net-based and Transformer-based diffusion models, including the Stable Diffusion series and FLUX series and adapts self-cross guidance as an effective reward for RL-based post-training and show improved subject diversity with no computational overhead at infe...
Weimin Qiu, Jieke Wang, Zhining Gu et al.· International Journal of Com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.