Skip to content

Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis

Jul 2026 · arXiv.org · Vol abs/2607.22147 · 0 citations · 42 references
Computer Science

TL;DR

This work proposes a visual relocalization method that departs from classical correspondence-based pipelines by directly estimating camera poses against a differentiable map representation built with 3D Gaussian Splatting (3DGS), and shows substantial gains in relocalization accuracy under challenging conditions.

Abstract

Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and sparse, weakly overlapping viewpoints induced by forward rover motion and unconstrained driving directions. Under these conditions, state-of-the-art image-to-image and image-to-map matching pipelines suffer significant performance degradation. In this work, we propose a visual relocalization method that departs from classical correspondence-based pipelines by directly estimating camera poses against a differentiable map representation built with 3D Gaussian Splatting (3DGS). Our key contribution is a geometry-aware training strategy that combines photometric and geometric losses, where the geometric supervision is provided for the first time by combining multi-view stereo (MVS) and LiDAR depths. We show that this joint optimization produces a 3DGS model that better fits the underlying scene geometry, leading to improved photometric and geometric consistency and more robust, accurate single-image 6-DoF pose estimation. Extensive experiments on data acquired in planetary-analog environments validate the effectiveness of our approach, showing substantial gains in relocalization accuracy under challenging conditions. Code is available at https://github.com/DLR-RM/multimodal-gsplat-relocalization.

View source

Similar papers

Open access Aug 2026

Using textureless, low-detailed 3D city models for visual localization

This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.

Yasmin Loeper, Markus Gerke, P. Fanta-Jende · 0 citations
Open access Aug 2026

Dynamic Neural Scene Reconstruction from Sparse Multi-View Observations with Geometry-Aware Depth and Correspondence Constraints

A novel framework designed to achieve high-fidelity dynamic neural scene reconstruction from highly sparse viewpoints by integrating geometry-aware depth priors and robust multi-view correspondence constraints is proposed, offering a scalable solution for dynamic scene capture.

Camila Wilson, D. J. Ross · 0 citations
Jul 2026

MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction

Results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.

Jinqian Yang, Yichen Wu, Wanhua Li et al. · 1 citation
Open access Aug 2026

Semantic-guided 3D Gaussian splatting for sparse-view reconstruction in industrial digital twins

A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.

Boyang Li, Tian-Han Gao, Zuan Gu et al. · 0 citations
Jul 2026

Enhance 3D Gaussian splatting for dynamic scenes: integrating semantic and geometric consistency

The proposed approach achieves competitive or superior performance compared with 3DGS, SpotLessSplats, T-3DGS, and RobustSplat on standard image-quality metrics, including peak signal-to-noise ratio, structural similarity index measure, and learned perceptual image patch similarity.

Wen Zheng, Guo Bao, Wenda Wang et al. · 0 citations
Jul 2026

MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors

This work proposes three learned matching heads: a LightGlue-style attention head with DoubleSoftmax scoring on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from the VGGT foundation model before pooling; and a multi-view extension that performs joint self-attention over segments drawn from several views at once, recovering transitive correspondences that strictly pairwise matchers cannot reach.

Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.