Towards a Modular Bin-picking Framework for Handling Object Pose Uncertainties
Frederik Hagelskjær
Jul 2026· International Conference on Automation, Control and Robotics Engineering· Vol abs/2607.13698· 0 citations· 36 references
Computer Science
TL;DR
This is the first framework that jointly addresses both grasping and object pose uncertainties using interchangeable modules, and there is ample opportunity to integrate additional modules, resulting in improved performance and flexibility.
Abstract
In recent years, there has been growing interest in robust robotic systems for precise bin-picking applications. To achieve reliable performance, such systems must address errors arising from both the object pose estimation and the grasping process. Although various approaches have been proposed, they typically target specific challenges and do not offer general solutions. In this paper, we present a modular framework that jointly handles both error types. The framework incorporates object pose distribution estimation to account for pose uncertainty, which frequently arises in situations with ambiguous observations where a single correct pose cannot be determined. To further reduce uncertainty, we introduce a second-viewpoint module that computes complementary pose distributions, which are subsequently fused. This fusion decreases overall uncertainty and improves system efficiency. Additionally, two independent modules are included to compensate for grasping errors. The modular design allows the components to be combined for optimal performance or used individually, depending on the physical setup. The proposed method is evaluated in a real-world setup with three different objects, with no errors, and all modules are shown to improve efficiency. These results suggest that incorporating pose distributions with grasping pose errors is a promising direction for developing more flexible and reliable robotic production systems. To the best of our knowledge, this is the first framework that jointly addresses both grasping and object pose uncertainties using interchangeable modules. We believe there is ample opportunity to integrate additional modules, resulting in improved performance and flexibility. The current framework is limited to pose uncertainties in SO(2), but it could be extended to SE(3), enabling additional modules to improve the system.
An innovative approach is introduced for the random bin-picking of planar objects by developing a multi-task model for instance segmentation and keypoint detection in 2D images and a grasp candidate selection strategy is proposed to enable reliable grasping in cluttered industrial environments.
The-Thinh Pham, Tuan-Khanh Nguyen, Chi-Cuong Tran et al.· Journal of Technical Educati...· 0 citations
A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.
Hui Zhang, Yue Wang, Kang An et al.· Signal, Image and Video Proc...· 0 citations
Industrial disassembly processes require robust and efficient perception systems capable of handling heterogeneous objects under real-world constraints. In automotive manufacturing and disassembly scenarios, components exhibit significant variability in geometry, size, symmetry, and placement, making it impractical to rely on a single, uniform object pose estimation strategy. At the same time, such environments impose strict requirements on the perception systems in terms of reliability, computational efficiency, and ease of deployment.This work presents an instance-based object perception pipeline with task-driven object pose estimation, developed within the SOPRANO EU project, for automated automotive door disassembly. The pipeline assumes a known set of object instances and depending on the objects’ geometric characteristics, the manipulation task requirements , the system performs (i) full 6D pose estimation, (ii) planar-constrained 3D localization via RGB-D lifting, or (iii) planar localization with structured multi-instance refinement and instance identification.The perception process consists of two stages: pose formulation selection and execution of the corresponding estimation method. Model-based 6D pose estimation using RGB data is applied to rigid objects, while RGB-D-based lifting of 2D detections is used for planar or quasi-planar elements. For structured arrangements of multiple instances of a single object type, such as screw arrays, a multi-instance matching strategy ensures consistent indexing and reduces ambiguity.The system is deployed as a modular perception service and validated in an automotive disassembly pilot. Experimental results demonstrate high accuracy across heterogeneous tasks, highlighting the benefits of aligning perception outputs with task-specific requirements.
Evangelos G. Sartinas, Athina Zacharia, Maria Pateraki· AHFE International· 0 citations
Estimation of the absolute pose of an object is an essential task for various robotic applications. Recently, incorporating gravity direction as prior information has emerged as a popular approach to simplify absolute pose estimation. However, developing a robust and efficient algorithm to solve this challenging problem remains a difficult question due to large amounts of mismatches. In addition, obtaining an accurate pose solution from selected inlier correspondences with gravity prior is still a research gap. In this paper, we propose a novel transformation strategy that exploits geometric relations derived from the gravity prior. Through transformation decoupling, the original 6 degrees of freedom (DoF) absolute pose estimation problem is simplified into a 4-DoFs problem: 1-DoF for the rotation angle and 3-DoFs for translation, significantly improving the efficiency. For the 1-DoF rotation angle, we apply a one-dimensional global voting algorithm for optimal estimation. Once the optimal rotation is obtained, the mismatched correspondences are preliminarily filtered, and translation estimation, a linear problem, can be easily solved. Furthermore, to obtain accurate pose results, we introduce a novel pose refinement algorithm to enhance the accuracy of both rotation and translation. Extensive experiments on synthetic data and three publicly available real-world datasets (TUM RGB-D, ETH3D, and RobotCar) demonstrate that the proposed method achieves stronger performance compared to existing state-of-the-art (SOTA) approaches. To further validate our method, we integrated it into ORB-SLAM2. The results on the KITTI dataset show it effectively reduces drift and improves trajectory alignment during relocalization. The source code will be released upon acceptance.
Hu Cao, Qian-Yi Yang, Xinyi Li et al.· 0 citations
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations
Robotic bin-picking of disordered, randomly stacked workpieces remains challenging because reliable grasping depends on an accurate estimate of object pose, yet many established solutions require high-precision 3D sensing, detailed object models, or large annotated datasets that raise the cost and effort of deployment on a new production line. This work presents a complete binocular vision framework that estimates workpiece pose by image matching and executes vision-guided grasping on a 6-DOF manipulator. A pose-annotated multi-view template library is constructed automatically through robot-driven image acquisition and compressed by a coarse-to-fine clustering scheme, and object pose is estimated by discriminative template matching with rigid refinement. To characterize the geometric reliability of the matched poses, an offline cross-modal analysis relates the 2D templates to a 3D reference model of the object and measures their agreement through region and contour reprojection metrics. Grasp configurations are then generated under orientation and collision constraints and corrected online by closed-loop visual feedback. Experiments on two representative workpieces show template-matching accuracy of 89–90% against classical and learned similarity measures, and grasp success between 81 and 87% across single-object and mixed scenes, outperforming the GraspNet baseline under the tested conditions. The framework offers an accurate and deployment-friendly route to robotic bin-picking.
Abdulrahman Usman Wunti, Ling-Xin Yu, Guangwei Li et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.