UniWorld-View is introduced, a unified framework for controllable large-baseline novel view synthesis from monocular inputs that integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation.
Abstract
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.
Novel view synthesis from sparse inputs requires both geometric grounding from the observed views and generative priors of unobserved regions, motivating recent hybrid methods that combine reconstruction and generation. However, existing methods bridge the two with rendered images or explicit 3D representations such as point maps or 3D Gaussians. Generation is thus conditioned on a lossy and imperfect projection of the scene, inheriting its errors, and reconstruction receives no signal from generation to correct them. We present RoGe, an end-to-end unified reconstruction and generation framework that removes this explicit bridge. It targets roaming within a scene anchored by sparse views: given a few posed images and a camera trajectory, it synthesizes a temporally coherent video along that trajectory. From the sparse input views, RoGe builds an implicit scene representation with a feed-forward reconstruction model, and queries it with target camera rays to obtain per-view geometric features. These features are injected into a video diffusion model as conditioning, without any 3D intermediate. Both modules are trained jointly, so the generation objective directly shapes its own geometric conditioning. We conduct experiments on DL3DV, where RoGe outperforms reconstruction-based, generation-based, and hybrid baselines on image-level metrics and video-level temporal consistency. Ablations confirm that ray-queried implicit features outperform both raw reconstruction tokens and rendered RGB as conditioning, and that joint training brings further gains. Our project page is at https://jerry-locker.github.io/roge/.
Xiaolei Lang, Ze Kang, Zehao Huang et al.· 0 citations
Novel view synthesis methods, such as neural radiance fields and 3D Gaussian splatting, offer a promising solution for photorealistic rendering. However, they remain challenged in few-shot settings, where models tend to overfit the limited supervised views, leading to artifacts such as quality fluctuations, degradation in distant views, and geometric inconsistencies. To address these issues, we introduce Human Perceptual Preference Optimization (HuPPO), a framework that incorporates human perceptual guidance into model training. HuPPO mitigates distortions by regularizing training dynamics with perceptual preference cues, thereby reducing the reliance on extensive supervised views. Specifically, HuPPO leverages human perception to identify and select candidate novel views, and introduces a corresponding objective function that steers optimization toward perceptually preferred outcomes. In addition, a meta-learning pipeline is integrated to promote the learning of generalizable representations. The framework is flexible and can be seamlessly applied to a wide range of neural rendering models without incurring additional inference overhead. Extensive experiments and analyses demonstrate that HuPPO achieves consistent improvements over state-of-the-art baselines.
Xiaoyu Xu, Jiebin Yan, Sheyang Tang et al.· IEEE Transactions on Visuali...· 0 citations
GenRec is introduced, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow, and attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones.
Ata Çelen, Jaewoo Jung, Federico Tombari et al.· 0 citations
This work presents FixAnything, a single model for fixing a wide range of rendering artifacts by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning.
Khiem Vuong, D. Ramanan, Srinivasa Narasimhan· 0 citations
DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion student.
This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.
Xinhui Liu, Can Wang, Wei Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.