Skip to content

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

Jul 2026 · arXiv.org · Vol abs/2607.16015 · 0 citations · 29 references
Computer Science

TL;DR

PIXIE is a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model, inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object.

Abstract

6D pose estimation remains a key challenge in robotics and computer vision, particularly in industrial environments. The deployment of currently available data-driven methods is often limited by resource-intensive data pipelines, reliance on textured 3D models, and sensitivity to geometric deviations caused by damages or assembly defects. We present PIXIE, a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model. Synthetic depth and normal maps are rendered from sampled reference viewpoints and matched to the query image via a pretrained cross-modality feature matcher. Matched keypoints are back-projected to obtain 2D--3D correspondences for PnP-based pose estimation. Relying exclusively on geometry makes the method inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object. We evaluate on widely-used public benchmarks, reporting state-of-the-art results on texture-less objects without object-specific training, and introduce a novel dataset with assembly defects, texture variations, and occlusion to demonstrate real-world applicability.

View source

Similar papers

Aug 2026

Accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios

A two-stage method for accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios with YOLO-based detection, Euclidean clustering, and moving least squares smoothing combined to extract high-quality target point clouds from noisy RGB-D observations is proposed.

Zhi-Xiang Zhang, Feng Luan, Wen-Hao Li · 0 citations
Preprint Sep 2026

Generalizable 6D Pose Estimation of Textureless Objects with Planar-based Gaussian Splatting

Estimating the 6D pose of textureless objects without prior CAD models remains a critical challenge due to the lack of appearance features. While recent generalizable approaches alleviate the dependence on object-specific models, their performance on low-texture objects is often limited by insufficient geometric constraints in the underlying representations. In this work, we propose PG-Pose, a geometry-aware framework combining Planar-based Gaussian Splatting (PGS) reconstruction and Geometry-driven pose optimization. In the offline representation extraction stage, three distinct representations of the object are extracted from multi-view reference RGB images with known poses. PG-Pose reconstructs a 3D Gaussian representation and renders high-fidelity depth maps to generate 3D point clouds through back projection. In the online pose inference stage, the initial pose of the input image is estimated by 2D-3D correspondence matching between the input image and the reconstructed 3D point clouds, followed by a PGS-Refiner for iterative pose optimization. Evaluations on the OnePose-LowTexture datasets, PG-Pose achieves an average accuracy of 94.2% ADD(S)@0.1d, with a 2.1% improvement average accuracy compared with the state-of-the-art (SOTA) GS-based approach. To further demonstrate the effectiveness of PG-Pose for industrial robots in grasping tasks, we deploy it on a dual-arm industrial robot and successfully realize the grasping task on an unseen object.

Jie Lu, Heng-Tan Zhang, Li Gong et al. · 0 citations
Open access Aug 2026

Using textureless, low-detailed 3D city models for visual localization

This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.

Yasmin Loeper, Markus Gerke, P. Fanta-Jende · 0 citations
Preprint Aug 2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.

Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al. · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations
Open access Sep 2026

Hybrid representation and adaptive multi-feature fusion for monocular 6D pose estimation in industrial assembly

Vision-based 6D pose estimation is critical for augmented reality-assisted assembly, human–robot collaboration, and quality inspection in intelligent manufacturing. However, performance degrades severely in complex industrial scenarios due to occlusion, varying lighting, textureless surfaces, and reflective parts. This work presents a monocular 6D pose estimation approach using hybrid representations and adaptive multi-feature fusion to address these challenges. A hybrid representation learning framework is designed to jointly predict keypoint heatmaps, relational vectors, semantic edges, masks, and visibility, thereby enhancing feature robustness. A multi-feature adaptive fusion strategy optimizes the pose by combining semantic and fine-grained general features. A structure-constrained correction module refines multi-object poses using assembly consistency constraints. Experiments on a custom industrial assembly dataset and the public Mono6D dataset show that the proposed method achieves 87.46% ADD (0.1d) and 86.82% 5 cm/5∘\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^\circ $$\end{document} accuracy, outperforming state-of-the-art methods. The custom dataset includes multiple weakly textured and reflective assembly parts under occlusion, lighting variation, and multi-viewpoint conditions. Furthermore, the complete system runs at approximately 18 FPS, with faster tracking once initialized. The approach supports reliable AR-assisted assembly and meets industrial deployment requirements. Our code and datasets are open-sourced at https://github.com/nengbinlv/HRMFPose, with the DOI: https://doi.org/https://doi.org/10.5281/zenodo.19574143.

Neng-Bin Lv, Zhang-Mao Xu, Yi Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.