Two-stage Monocular 6D Pose Estimation for Small Cubic Objects
Abstract
Monocular 6D object pose estimation is an important per- ception problem for robotic manipulation. For high-precision operations such as block grasping, alignment, and placement, the perception mod- ule must provide sufficiently accurate and stable pose estimates rather than coarse object localization alone. This requirement is particularly challenging for small cubic objects due to weak texture, limited visual cues, strong rotational symmetry, and sensitivity to ROI quality. In this paper, we study monocular 6D pose estimation of small cubic objects from a single RGB image and propose a two-stage manipulation- oriented framework. In the first stage, a detection-guided ROI-based re- gression model is used to estimate object pose under symmetry-aware supervision. In the second stage, we further explore a MuJoCo-based re- rendering strategy to construct consistency-enhanced training samples for refinement, aiming to improve adaptation to manipulation-relevant visual conditions. Experiments on a unified MuJoCo-generated dataset show that, on the main single-block benchmark, the proposed method achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics. Preliminary results in multi-block scenes further suggest favorable generalization to more complex visual condi- tions.