2026· IEEE Transactions on Instrumentation and Measurement· Vol 75, pp. 5016113-5016113· 0 citations· 52 references
Abstract
Visual localization is a key technology in many vision-based measurement applications, aiming to estimate the camera pose of a query image in a known environment. However, most existing methods rely on heavy scene-specific representations, such as explicit 3-D map construction or per-scene training. Constructing and maintaining such representations introduces nonnegligible computational overhead, storage burden, and long-term maintenance costs. To address this issue, we propose a novel visual localization pipeline that uses a set of posed reference images as a lightweight scene representation and localizes query images without explicit 3-D map construction or scene-specific training. Specifically, we exploit a geometric foundation model to infer local multiview geometry from the query image and its retrieved references. Since the predicted geometry is expressed in an arbitrary local coordinate system with unknown scale, a key challenge is how to recover an accurate metric pose of the query image from such local predictions. To address this challenge, we design a global pose recovery strategy that first registers the predicted local geometry to the world coordinate system through joint center–orientation similarity alignment using the posed reference images as global anchors, and then refines the query pose by optimizing query-associated 3-D landmarks under multiview 2-D–3-D geometric constraints. The experimental results on multiple benchmark datasets show that our method achieves competitive localization performance and improved robustness under sparse reference-view settings and challenging viewpoint or appearance variations, reducing the average translation and rotation errors of the strongest Unseen baseline from 55 cm and 0.56° to 13 cm and 0.23° on Cambridge Landmarks, respectively.
This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.
Yasmin Loeper, Markus Gerke, P. Fanta-Jende· The International Archives o...· 0 citations
A scalable visual localization pipeline that combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation is introduced, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches.
Jonas Meyer, S. Nebiker, P. Theiler et al.· arXiv.org· 0 citations
UniQuery4R is presented, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention, and introduces a direction-magnitude parameterization of scene flow with separate supervision for moving and static points.
Tiancheng Chen, Sheng Tang, Wenhua Jin et al.· 0 citations
This work proposes a visual relocalization method that departs from classical correspondence-based pipelines by directly estimating camera poses against a differentiable map representation built with 3D Gaussian Splatting (3DGS), and shows substantial gains in relocalization accuracy under challenging conditions.
M. Peribañez, Javier Civera, Rudolph Triebel et al.· arXiv.org· 0 citations