RGA6D: Regional Geometry-Aware Correspondence Reasoning for Depth-Only Model-Free 6D Pose Estimation
Abstract
Estimating the 6D pose of unseen objects from depth observations is fundamental for robotic perception, particularly for texture-less objects where RGB appearance provides limited cues. Despite rich geometric information provided by depth images, depth-only model-free 6D pose estimation remains largely underexplored. Its core challenge lies in establishing reliable 3D–3D correspondences under noisy observations and structural ambiguities caused by repetitive structures and object symmetries. To address these problems, we propose RGA6D, a regional geometry-aware correspondence reasoning framework for depth-only model-free 6D pose estimation. Specifically, we first introduce a Transformer-based bottleneck architecture for regional context reasoning, enabling regional representatives to progressively aggregate, interact, and propagate stable structural cues to learn noise-robust geometric representations for geometrically consistent correspondence construction. We then hierarchically reason over regional correspondence consensus through region-wise geometric alignment and transformation-consistent grouping across regions, generating a compact set of reliable pose hypotheses to alleviate ambiguities. Finally, a high-resolution pose refinement module further improves geometric alignment. Extensive experiments on the challenging texture-less T-LESS and Linemod datasets show that RGA6D achieves outstanding performance among existing depth-only baselines while remaining highly competitive with RGB(-D) methods.