Skip to content

Geometry-Informed Auxiliary Priors for Consistent 3D Content Generation.

Aug 2026 · IEEE Transactions on Visualization and Computer Graphics · Vol PP · 0 citations
Medicine

TL;DR

This work proposes GeometryAP, a training-free framework for consistent 3D content generation via geometry-informed auxiliary priors in a two-stage "generate-then-reconstruct" pipeline, and proposes the 3D Gaussian-based Global-to-Local Mesh optimizer (GaussGLMeshizer).

Abstract

Despite recent advancements in single-view 3D re construction leveraging multi-view diffusion (MVD), they still suffer from multi-view inconsistency and low-fidelity reconstruction. Existing solutions often introduce priors through prolonged, large-scale training, which is costly, and they struggle with aligning features with conditioning signals. In this work, we propose GeometryAP, a training-free framework for consistent 3D content generation via geometry-informed auxiliary priors in a two-stage "generate-then-reconstruct" pipeline. Our approach consists of two critical components: (1) To enhance cross-view consistency, we design the Auxiliary Prior-guided Attention Mechanism (APAM), which uses the MVD's own pseudo-normal noise features as an auxiliary prior. APAM fuses this 3D prior information to activate crucial attention regions, ensuring the 3D consistency of generated multi-view images. (2) To achieve robust reconstruction from the generated multi view images, we propose the 3D Gaussian-based Global-to-Local Mesh optimizer (GaussGLMeshizer). It leverages 3DGS as an intermediate representation to generate normal maps from both auxiliary and basic views, which serve as auxiliary priors. The global optimizer uses normal maps from auxiliary views to ensure high-fidelity detail on unknown views, while the local optimizer refines existing details using basic view normals. Moreover, a top-k mechanism is applied at different stages to mitigate prior errors, further enhancing robustness. Experiments on single view inputs demonstrate that GeometryAP outperforms state of-the-art baselines in geometry consistency, detail integrity, and structure fidelity without data training.

View source

Similar papers

Preprint Aug 2026

Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction

This work investigates the integration of geometric priors, in the form of predicted normal and depth maps, into the 3DGS framework to improve the reconstruction quality and reveals that multi-view predictions, as they are done by the recent visual geometry grounded transformer (VGGT), outperform single-view alternatives.

Hongyu Zhou, Zorah Lähner · 0 citations
Jul 2026

Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

Axolotl3D is presented, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud that synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning.

A. Hu, Maria Shugrina · 1 citation
Preprint Sep 2026

DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting

While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction and novel-view synthesis, scenarios with limited input views often lead to poor reconstruction quality and artifacts in rendered novel views. Recent efforts attempt to utilize powerful diffusion priors, yet they typically process rendered and reference views concatenated along an additional dimension in a single network. These methods overlook an inherent nature that different views should maintain appearance similarity but differ in structure due to view shifts, leading to blur caused by conflicts between the two properties. In this paper, we propose DualDiff, a novel pipeline that leverages dual diffusion priors with a Structure-Appearance Attention (SAA) module to introduce reference guidance for refining low-quality novel views rendered from flawed 3D representations. Specifically, we retain one diffusion branch to focus on extracting structural information from the low-quality novel views, while introducing another branch to ensure appearance consistency with reference views. Furthermore, we present a 3D reconstruction framework named DualDiff3D, which integrates a reliability-enhanced Render-Refine-Optimize (RRO) loop to progressively and robustly incorporate the refined novel views, yielding more accurate 3DGS. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods even in the inference-only setting, with further performance gains achievable through training. Our code and pre-trained weights are available at https://github.com/Akaneqwq/DualDiff3D.

Qian Wang, Yu Wang, Weiqi Li et al. · 0 citations
Preprint Aug 2026

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.

Xinhui Liu, Can Wang, Wei Jiang et al. · 0 citations
Open access Jul 2026

Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline

Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extreme lack of information. Recently, diffusion models pre-trained on large-scale datasets served as 2D priors are used to solve the ill-posed task but suffer from color deviation and view inconsistency, which can be curbed by using diffusion models fine-tuned with 3D annotated data served as 3D priors. However, 3D priors lack high-frequency details, which cannot be solved by direct complementation with 2D priors in spatial domain for introducing erroneous low-frequency 2D prior guidance. In this paper, we revisit the characteristics of different diffusion priors from the frequency perspective. Based on our observations, we theoretically present a unified framework of hybrid optimization using multiple diffusion priors in frequency domain. Under this framework, we further propose Morpheus3D, a pipeline of 3D object generation from any single unposed image in the wild. Morpheus3D enhances 3D prior with high-pass image-prompt 2D prior guidance to reconstruct high-quality 3D objects while effectively suppressing view inconsistency, low-frequency color deviation, and high-frequency lacking problems. Both quantitative and qualitative experiments on the public and our collected datasets with complex textures show that our method exhibits significant improvements in generation quality.

Qisen Wang, Yifan Zhao, Jia Li · 0 citations
Open access Oct 2025

SaLon3R: Structure-Aware Long-Term Feedforward 3D Reconstruction from Unposed Images

This work proposes SaLon3R, a novel framework for Structure-aware, Long-term 3DGS Reconstruction that effectively prunes the redundant 3DGS and resolves artifacts in a single feed-forward pass, and introduces a 3D Point Transformer to overcome geometric inconsistencies caused by long-term accumulative errors.

Jiaxin Guo, Tongfan Guan, Wen-Zhen Dong et al. · 5 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.