DAP-Pose accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment, and incorporates physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift.
Abstract
Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.
Accurate monocular pose estimation of spacecraft is essential for on-orbit servicing and proximity operations, yet this remains challenging because of large variations in the target scale and limited onboard computational resources. To address these challenges, this paper proposes SAPose, a lightweight geometry-guided...
DOU-Pose is proposed, a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline to enhance the discriminative capability of scene coordinate regression through improved feature extraction and replaces standard convolutional layers with Depthwise Over-parameterized Convolution (D...
Xin'an Qiu, Li-Wen Wang, Zezheng Dong et al.· Italian National Conference...· 0 citations
GS-CPE (Gaussian Splatting based Camera Pose Estimation), a coarse-to-fine framework for 6-DoF camera pose estimation that unifies geometry-based coarse pose estimation with robust 3D Gaussian Splatting based pose refinement, is introduced.
Real-time 6DoF object pose estimation on resource-constrained hardware remains challenging, as accurate correspondence-based and refinement pipelines typically rely on non-differentiable PnP/RANSAC stages or costly iterative refinement, while recent foundation-model-based approaches incur inference costs that are prohi...
P. Kühn, D. Nguyen, Saptarshi Neil Sinha et al.· 0 citations
We propose Sen-Cap, a Sensor-Flexible and Noise-Resilient 3D human motion Capture framework that integrates multi-modal data from LiDAR and camera. While multi-modal sensors provide richer information than single-modal sensors, existing approaches still suffer from two core challenges. First, multi-modal alignment/matc...
Aoru Xue, Yujing Sun, Yiming Ren et al.· 0 citations
AeroMotion6D is proposed, a temporal transformer-based framework for monocular UAV 6D pose estimation from RGB video that consists of an adaptive context fusion mechanism that can incorporate past context information into the current estimation process and a persistent pose memory module that can convey pose-related in...
Mohammad Al Qaderi, M. Hayajneh, Alaa Alghazo et al.· Robotics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.