Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extreme lack of information. Recently, diffusion models pre-trained on large-scale datasets served as 2D priors are used to solve the ill-posed task but suffer from color deviation and view inconsistency, which can be curbed by using diffusion models fine-tuned with 3D annotated data served as 3D priors. However, 3D priors lack high-frequency details, which cannot be solved by direct complementation with 2D priors in spatial domain for introducing erroneous low-frequency 2D prior guidance. In this paper, we revisit the characteristics of different diffusion priors from the frequency perspective. Based on our observations, we theoretically present a unified framework of hybrid optimization using multiple diffusion priors in frequency domain. Under this framework, we further propose Morpheus3D, a pipeline of 3D object generation from any single unposed image in the wild. Morpheus3D enhances 3D prior with high-pass image-prompt 2D prior guidance to reconstruct high-quality 3D objects while effectively suppressing view inconsistency, low-frequency color deviation, and high-frequency lacking problems. Both quantitative and qualitative experiments on the public and our collected datasets with complex textures show that our method exhibits significant improvements in generation quality.
Qisen Wang, Yifan Zhao, Jia Li· International Journal of Com...· 0 citations
Neural Radiance Fields can achieve photo-realistic rendering results, but the occlusion in front of the target object is a common and extreme scenario in practice that cannot be neglected. The prevailing works attempt to remove the occlusions using external 2D visual priors, which are not constrained to provide 3D-consistent guidance for the specific scenarios. In this paper, we propose UncNeRF, which utilizes multi-view clues from captured defective images to uncover the heavily occluded object. Specifically, we provide additional multi-view complementary optimization supervisions using object-centric forward warping and enhance the target object reconstruction by sampling pseudo-training views and introducing external spatial-relation regularization. To evaluate the reconstruction performance of occluded objects, we present the challenging and diverse Heavy Occlusion Removal (HOR) dataset consisting of synthetic and real-world scenes, whose target objects to be reconstructed are heavily occluded. Experimental results show that our method achieves state-of-the-art performance in heavy occlusion removal compared to other methods.
Jiawei Ma, Qisen Wang, Yifan Zhao et al.· IEEE Transactions on Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.