Skip to content

DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

Jul 2026 · arXiv.org · Vol abs/2607.23755 · 0 citations · 34 references
Computer Science

TL;DR

DAP-Pose accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment, and incorporates physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift.

Abstract

Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.

View source

Similar papers

Open access Sep 2026

Lightweight Monocular Relative Pose Estimation of Spacecraft via Scale-Adaptive Multi-Task Learning

Accurate monocular pose estimation of spacecraft is essential for on-orbit servicing and proximity operations, yet this remains challenging because of large variations in the target scale and limited onboard computational resources. To address these challenges, this paper proposes SAPose, a lightweight geometry-guided...

Zhi-Wei Hu, Yi-Jie Zhang, Bo-Wen Hou et al. · 0 citations
Open access Aug 2026

DOU-Pose: Robust Camera-Based Visual Localization for Autonomous Vehicles in Repetitive and Low-Texture Intelligent Transportation Environments

DOU-Pose is proposed, a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline to enhance the discriminative capability of scene coordinate regression through improved feature extraction and replaces standard convolutional layers with Depthwise Over-parameterized Convolution (D...

Xin'an Qiu, Li-Wen Wang, Zezheng Dong et al. · 0 citations
Preprint Aug 2026

GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting

GS-CPE (Gaussian Splatting based Camera Pose Estimation), a coarse-to-fine framework for 6-DoF camera pose estimation that unifies geometry-based coarse pose estimation with robust 3D Gaussian Splatting based pose refinement, is introduced.

Huai-Yuan Weng, C. Yeum, Su-Min Kang · 0 citations
Preprint Aug 2026

TinyDETR-Pose: Towards End-to-End Real-Time Single-Stage 6DoF Object Pose Estimation with Lightweight Transformers

Real-time 6DoF object pose estimation on resource-constrained hardware remains challenging, as accurate correspondence-based and refinement pipelines typically rely on non-differentiable PnP/RANSAC stages or costly iterative refinement, while recent foundation-model-based approaches incur inference costs that are prohi...

P. Kühn, D. Nguyen, Saptarshi Neil Sinha et al. · 0 citations
Preprint Aug 2026

Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration

We propose Sen-Cap, a Sensor-Flexible and Noise-Resilient 3D human motion Capture framework that integrates multi-modal data from LiDAR and camera. While multi-modal sensors provide richer information than single-modal sensors, existing approaches still suffer from two core challenges. First, multi-modal alignment/matc...

Aoru Xue, Yujing Sun, Yiming Ren et al. · 0 citations
Open access Jul 2026

Real-Time Temporally Consistent Monocular 6D UAV Pose Estimation for Onboard Aerial Perception

AeroMotion6D is proposed, a temporal transformer-based framework for monocular UAV 6D pose estimation from RGB video that consists of an adaptive context fusion mechanism that can incorporate past context information into the current estimation process and a persistent pose memory module that can convey pose-related in...

Mohammad Al Qaderi, M. Hayajneh, Alaa Alghazo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.