Skip to content
Open access

DOU-Pose: Robust Camera-Based Visual Localization for Autonomous Vehicles in Repetitive and Low-Texture Intelligent Transportation Environments

Aug 2026 · Italian National Conference on Sensors · Vol 26 · 0 citations · 52 references
Medicine

TL;DR

DOU-Pose is proposed, a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline to enhance the discriminative capability of scene coordinate regression through improved feature extraction and replaces standard convolutional layers with Depthwise Over-parameterized Convolution (DO-Conv).

Abstract

Accurate and robust vehicle localization is essential for autonomous driving. However, existing visual pose estimation methods often struggle in scenarios dominated by repetitive structures or sparse textures. These conditions lead to ambiguous predictions of 3D scene coordinates and a high proportion of structured outliers—erroneous predictions forming coherent clusters that deceive standard estimators. To address these limitations, this paper proposes DOU-Pose (Depthwise Over-parameterized U-shaped Pose estimation), a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline. The core idea is to enhance the discriminative capability of scene coordinate regression through improved feature extraction. Specifically, we replace standard convolutional layers with Depthwise Over-parameterized Convolution (DO-Conv), which introduces auxiliary learnable depthwise kernels during training to enrich the representational capacity of the network, while allowing their fusion into a single kernel for inference. Furthermore, a U-shaped regression network with transposed convolutions is designed to preserve spatial details and strengthen fine-grained geometric reasoning. The entire pipeline is trained end-to-end by coupling dense scene coordinate prediction with a differentiable robust estimator. Extensive experiments demonstrate that DOU-Pose achieves competitive performance on public benchmarks and clear robustness improvements on the self-collected Campus-AV dataset, especially in repetitive and low-texture outdoor driving scenarios.

Read PDF

Similar papers

Preprint Aug 2026

Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models

Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the first two requirements, but its low metric accuracy makes it less suitable than standard VL approaches based on local features or neural representations. This limitation can be overcome by integrating VPR with the accurate local trajectory estimates produced by feed-forward neural 3D geometry (FF3D) models. In this paper, we address sequential appearance-based localization through a topometric framework that iteratively combines probabilistic VPR with FF3D metric pose estimation in controlled image sets. Our approach proposes an automatic offline mapping tool that models the topometric pose-appearance interaction in the different parts of the scene. This map is later employed by an online particle filter that estimates the pose from odometry and belief over places for FF3D inference, successfully incorporating neural metric estimation into probabilistic appearance-based localization. We extensively evaluate the framework on three known benchmarks, demonstrating substantial improvements over existing appearance-based methods. The modularity of our approach allows the descriptor extractor and FF3D model to remain interchangeable, and a focused analysis further shows that sequential belief can mitigate severe failures under perceptual aliasing.

Eulogio Quemada-Torres, Alberto Jaenal, Francisco-Angel Moreno et al. · 0 citations
Review Jul 2026

Accuracy potential of visual localization exploiting high-end street-level imagery

A scalable visual localization pipeline that combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation is introduced, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches.

Jonas Meyer, S. Nebiker, P. Theiler et al. · 0 citations
Open access Aug 2026

Using textureless, low-detailed 3D city models for visual localization

This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.

Yasmin Loeper, Markus Gerke, P. Fanta-Jende · 0 citations
Jul 2026

RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection

In intelligent transportation systems, roadside 3D object detection provides wide-area perception crucial for traffic understanding, cooperative early warning, and safe autonomous driving. However, existing methods suffer from high sensitivity to camera extrinsics; even slight deviations (whether manifesting as transient jitter or persistent drift) can be significantly amplified by projective geometry. This cascade results in severe feature misalignment and degraded localization. To mitigate this limitation, we propose RECO, a region-aware extrinsic compensation framework that corrects extrinsics using piecewise 6-DoF pose offsets. RECO predicts a learnable range boundary to partition the scene into near and far regions, estimating region-specific pose corrections. A differentiable sigmoid gate then smoothly blends the two compensated geometries to preserve continuous BEV sampling and facilitate stable optimization. To supervise the refinement of extrinsics, we introduce an auxiliary reprojection loss that compares 2D bounding boxes projected from 3D ground truth against 2D annotations, optimizing it jointly with the standard detection objective. Extensive experiments on the DAIR-V2X-I and Rope3D benchmarks under extrinsic perturbations demonstrate consistent improvements over state-of-the-art baselines across both yaw and $z$-axis deviations. RECO also generalizes from transient perturbations to persistent shifts, maintaining highly competitive performance under strict calibration uncertainty.

Junsheng Du, Zhaocheng He, Yuhuan Lu · 0 citations
Jul 2026

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

PIXIE is a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model, inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object.

Leon Jungemeyer, A. Magaña, Gautham Mohan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.