Skip to content
Open access

BEV-LOC: Real-Time and Lightweight Cross-View Localization via Online BEV Mapping

Jul 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XLIX-B2-2026, pp. 715-720 · 0 citations · 17 references

TL;DR

BEV-LOC is presented, a lightweight and real-time cross-view geolocalization method that employs Bird’s Eye View encoder that learns to transform 360-degree multi-PV images into a local high-definition (HD) BEV map and is performed using Intersection Over Union (IoU)-based template matching with an offline global map.

Abstract

Abstract. This paper presents a deep learning and classical computer vision framework for cross-view geolocalization using 360-degree multi-perspective view (PV) images and an offline global map. Recent studies on cross-view geolocalization typically rely on deep learning models to localize panoramic PV images by matching them with reference satellite imagery. However, such approaches face practical limitations in real-world deployments, due to their dependence on large-scale GPU resources and the need to store extensive satellite image datasets. To address these challenges, we propose BEV-LOC, a lightweight and real-time cross-view geolocalization method. BEV-LOC employs Bird’s Eye View (BEV) encoder that learns to transform 360-degree multi-PV images into a local high-definition (HD) BEV map. The localization is then performed using Intersection Over Union (IoU)-based template matching with an offline global map. Our architecture achieves real-time performance at 30 FPS without the need for high-end GPU hardware and delivers a high positioning accuracy of 1.2 meters.

Read PDF

Similar papers

Preprint Aug 2026

DPA-I2P: Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration in Autonomous Driving

Image-to-Point Cloud Registration aims to estimate the camera pose of a given image within a 3D scene point cloud, which is a fundamental task in autonomous driving and large-scale outdoor localization. Recent implicit correspondence learning methods have improved registration performance by learning cross-modal alignment in an end-to-end framework, leading to more accurate camera pose estimation. However, due to the inherent modality discrepancy between images and sparse LiDAR point clouds, reliable cross-modal correspondence learning remains challenging. To address this issue, we propose Depth-Guided Projective Alignment for Image-to-Point-Cloud Registration (DPA-I2P). Unlike naive depth or feature concatenation, Ray-Conditioned Metric Depth Encoding (RMDE) and Projection-Consistent Vision Lifting (PVL) exploit depth and visual cues in a structured, geometry-aware manner. In addition, Cross-Modal Query Pruning (CQP) suppresses unreliable queries during early refinement to improve matching stability. Experiments on KITTI and nuScenes demonstrate the effectiveness of the proposed method. On KITTI, DPA-I2P reduces RTE and RRE by 45.0% and 55.6% over the strongest implicit baseline, respectively. On nuScenes, DPA-I2P also improves registration accuracy over the evaluated baselines, suggesting better transferability to different driving scenes.

Wen-Xin Zhang, Hang Li, Zhiwei Xu et al. · 0 citations
Preprint Sep 2026

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR maps, even under severe seasonal, weather, and illumination changes. The key idea is to convert both camera images and LiDAR scans into a unified depth image representation so that a single VFM backbone with an aggregation head can learn modality-invariant global descriptors. To make pairwise metric learning faithful to scene geometry, we introduce a geometry-aware overlap miner: after cross-modal scale alignment of camera and LiDAR depth, we forward-warp measurements between views to compute a pixel-level overlap score. This score relabels ambiguous pairs and adaptively modulates the positive margin in a multi-similarity loss to avoid overfitting on weakly overlapping views. Extensive experiments on KITTI odometry and Boreas demonstrate strong performance and robustness across seasons, weather, and day/night. On KITTI, DXPR achieves near-perfect Recall@1 on most sequences and outperforms prior CMPR baselines. On Boreas, DXPR achieves intra-sequence performance on par with a strong single-modal baseline (DINOv2-SALAD), while showing clear improvements in the more challenging inter-sequence setting. Compared with RangeBEV, our method consistently performs better in both intra- and inter-sequence evaluations, demonstrating robustness under diverse seasonal and illumination changes.

Yu-Hang Han, Youngseok Jang, Seungwon Roh et al. · 0 citations
Preprint Aug 2026

Pixel-wise Geo-registration of Drone and Satellite Images

Pixel-level cross-view geo-registration aims to align a query image (e.g., drone) to a geo-referenced satellite map so that every query pixel can be mapped to real-world GPS coordinates. Despite strong progress in cross-view geo-localization, existing benchmarks largely provide only GPS labels, limiting evaluation to a single coordinate per image and leaving dense geodetic alignment underexplored. We introduce SkyReg, a dataset and standardized benchmark for pixel-level drone-to-satellite geo-registration, providing dense per-pixel geo-location supervision across diverse settings (orthographic and perspective), scene types (urban, landmark-centric, suburban/rural), and camera configurations. Using SkyReg, we evaluate a broad set of baselines spanning retrieval, feature matching, homography-based alignment, and feed-forward 3D reconstruction. Finally, cross-view pairs from SkyReg, we train a geometry-aware reconstruction pipeline that achieves state-of-the-art results,improving performance by a significant margin.

Qingyang Liu, D. Shatwell, P. Kulkarni et al. · 0 citations
2026

A Cross-View Geolocalization Method Based on Frequency-Spatial Feature Enhancement and Dynamic Margin Constraint

In global navigation satellite system (GNSS)-denied scenarios, cross-view geolocalization (CVGL) provides an effective solution for autonomous uncrewed aerial vehicle (UAV) localization. However, viewpoint and scale differences across platforms may cause variations in texture structures and spatial distributions for the same geographic region, making it more difficult to learn consistent and discriminative cross-view representations in CVGL. Meanwhile, visually similar but geographically distinct samples can reduce the separability between true matches and hard negatives. To address these issues, we propose the frequency-spatial feature enhancement and dynamic margin constraint (FSDC) network, which integrates the frequency-aware recalibrated spatial (FARS) module and the margin-based dynamic contrastive learning (MDCL) strategy. The FARS module enhances local structural representations through bidirectional complementary interaction between frequency-domain and spatial structural features, guiding the shared encoder to learn more consistent cross-view representations, while the MDCL strategy imposes adaptive bounded constraints on hard-negative samples to improve feature discriminability and training stability. Experimental results on three benchmarks show that FSDC provides a favorable tradeoff among cross-view retrieval accuracy, model complexity, and inference efficiency.

Weiquan Wang, Yanfei Peng, Lei Ma et al. · 0 citations
Conference Jul 2026

Low-Cost Urban Road Mapping Via Dual-View Fusion of Panoramic Images

High-definition (HD) road maps are critical for autonomous navigation and intelligent transportation systems. However, single front-view pipelines suffer from unilateral occlusions and a narrow field of view (FoV), whereas conventional multi-sensor bird's-eye-view (BEV) systems improve coverage at the cost of increased hardware requirements, calibration complexity, and computation. This work addresses the problem of achieving robust, wide-coverage HD mapping under urban occlusions using a single, low-cost panoramic camera. A lightweight dual-view fusion framework is introduced for incremental road mapping from panoramic images. The method introduces three technical contributions: (1) a single-sensor dual-view construction that extracts front and rear perspective views from one panoramic camera via FoV-aware projection; (2) a geometry-consistent BEV fusion module that integrates inverse perspective mapping (IPM), pose-stabilized stitching, and patch-level merging to suppress parallax and motion jitter while recovering markings occluded in one view but visible in the other; and (3) a lightweight incremental pipeline that reduces deployment and inter-sensor calibration overhead relative to ring-camera systems. Experiments on a self-built dataset of 80 test panoramas with five road-element classes under unilateral or moderate occlusion show that dual-view fusion improves marking completeness by 33.3%, reduces geometric deviation (PSC) by 26.7%, and improves shape regularity (RARC) by 7.7% over a front-view-only baseline. The results support panoramic dual-view fusion as a practical low-cost compromise between limited single-view coverage and high-complexity multi-sensor platforms.

Zhebin Zhao, Yaxin Li, Hongsheng Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.