Skip to content

Accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios

Aug 2026 · The Visual Computer · Vol 42 · 0 citations · 42 references

TL;DR

A two-stage method for accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios with YOLO-based detection, Euclidean clustering, and moving least squares smoothing combined to extract high-quality target point clouds from noisy RGB-D observations is proposed.

View source

Similar papers

Jul 2026

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

PIXIE is a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model, inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object.

Leon Jungemeyer, A. Magaña, Gautham Mohan et al. · 0 citations
Preprint Aug 2026

Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real

The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real. The first is a geometry-first pipeline: YOLO11x detection, homography lifting to the world frame, class-level 3D size priors, multi-camera fusion, world-coordinate tracking, and offline tracklet stitching. The second is estimated-depth pseudo-LiDAR: monocular depth (D4RT, Metric3D~v2) back-projected into a fused point cloud and passed to a 3D detector (V-DETR), mirroring prior point-cloud winners that used depth. The gap is decisive: geometry-first reaches 13.0 3D HOTA (51.6 LocA), whereas pseudo-LiDAR collapses to 0.12 (9.2 LocA). We trace the collapse to cross-view inconsistency of monocular depth---scale correction is necessary but not sufficient---which domain-adaptation fine-tuning does not repair within budget. Within the geometry pipeline, offline stitching is the only intervention that helps; SAHI detection, appearance Re-ID, learned lifting, RT-DETR ensembling, test-time augmentation, and domain randomization all fail to beat the baseline detector. The bottlenecks are complementary: detection quality bounds the geometry route (DetA), localization consistency bounds pseudo-LiDAR (LocA). We release a complete, reproducible RGB-only pipeline and ablation.

Abdullah Naeem, Anav Katwal, Ayon Dey et al. · 0 citations
Jul 2026

Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment

Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate system.However, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work proposes a unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models. The system integrates anchor-free detection (YOLOX), robust tracking (BoT-SORT), omni-scale appearance embedding (OsNet), pose estimation (HRNet via MMPose), and transformer-based geometric reconstruction using the Visual Geometry Grounded Transformer (VGGT). A pose-guided 3D lifting strategy projects head keypoints onto a reconstructed manifold, eliminating dependence on ground-plane homography. Global identity association is formulated as hierarchical agglomerative clustering under a joint appearance-geometry cost with strict velocity gating. Evaluation on the AI City Challenge 2024 demonstrates a HOTA score of 53.13% without access to ground-truth calibration matrices, establishing a strong baseline for purely vision-based 3D tracking.

Ponleur Veng, Dominique Vaufreydaz, Phutphalla Kong · 0 citations
Preprint Aug 2026

A Height-Constrained 2-Point Minimal Solver for Pose Estimation from Active LED Markers with Event Cameras

In many autonomous applications requiring real-time localization, active marker-based systems are preferred due to their low latency and ease of deployment compared to computationally demanding feature-based methods. Event~\mbox{cameras} offer high temporal resolution and minimal delay and are commonly used with active LED markers for robust real-time localization. Existing methods typically rely on Perspective-n-Point (PnP) solvers for pose estimation. However, structured marker layouts can be challenging to deploy in space-constrained scenarios, while partial self-motion information (e.g., gravity direction and altitude) is readily available from onboard sensors. We derive a robust and accurate minimal solver that estimates camera pose from only two LED markers by incorporating known tilt angle and camera height measured by an onboard sensor, such as an IMU or an altimeter. The proposed formulation uniquely determines the camera pose through both a closed-form and a linear least-squares solution. We further analyze degenerate configurations and characterize the conditions under which height information does not contribute to rotation estimation. For evaluation, we developed an event-based active marker system to collect real-world data with ground truth from a motion capture system. Experiments on both synthetic and real data demonstrate improved accuracy over the state-of-the-art P2P solver and competitive performance relative to P3P.

Runze Yuan, Alexander Kappler, J. Zhang et al. · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations
Open access Aug 2026

GeoHash3D: Robust and Efficient 6-DoF Pose Estimation from 3D Marker Sets

Accurately measuring the six degrees of freedom (6-DoF) pose of a sample is a critical prerequisite for applications in robotic sample handling, optical metrology, and industrial quality control. Our pose estimation problem requires matching a small, partial observation of 4–20 discrete 3D points to a reference point set of up to 50 points. Popular registration algorithms such as Go-ICP, TEASER++, and MAC are designed for large-scale correspondence problems and do not perform reliably in our operating regime. Therefore, we propose an improved geometric hashing method that robustly estimates the rigid transformation in the presence of noise and outliers, and demonstrate its effectiveness and speed using both simulated and real-world datasets.

Lijiu Wang, Kailas Mahalinga Upadhyaya, Oguz Kedilioglu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.