Skip to content

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

Aug 2026 · Signal, Image and Video Processing · Vol 20 · 0 citations · 43 references

TL;DR

A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.

View source

Similar papers

Open access Jul 2026

A Robust Visual Grasping Method for Robots in Cluttered and Stacked Scenes

An iterative closed-loop optimization framework that deeply couples SAM with FoundationPose and designs a multi-dimensional confidence assessment module that integrates both the 2D image domain and the 3D geometric domain to comprehensively evaluate the reliability of the current pose.

Zhiqiang Gao, Mengqi Li, Huihui Bai et al. · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

Monocular 6D object pose estimation is an important per- ception problem for robotic manipulation. For high-precision operations such as block grasping, alignment, and placement, the perception mod- ule must provide sufficiently accurate and stable pose estimates rather than coarse object localization alone. This requirement is particularly challenging for small cubic objects due to weak texture, limited visual cues, strong rotational symmetry, and sensitivity to ROI quality. In this paper, we study monocular 6D pose estimation of small cubic objects from a single RGB image and propose a two-stage manipulation- oriented framework. In the first stage, a detection-guided ROI-based re- gression model is used to estimate object pose under symmetry-aware supervision. In the second stage, we further explore a MuJoCo-based re- rendering strategy to construct consistency-enhanced training samples for refinement, aiming to improve adaptation to manipulation-relevant visual conditions. Experiments on a unified MuJoCo-generated dataset show that, on the main single-block benchmark, the proposed method achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics. Preliminary results in multi-block scenes further suggest favorable generalization to more complex visual condi- tions.

Xinmiao Du · 0 citations
#artificial intelligence Preprint Aug 2026

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

A reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN), is proposed, offering a scalable and adaptable solution for contact-rich manipulation tasks.

Amir Arsalan Nematollahi, Shayan Ahmadi, M. T. Masouleh et al. · 0 citations
Open access 2026

Automated Model Selection for Task-Specific RGB–Tactile Fusion in In-Hand Grasp Pose Estimation

Robotic grasping of small, low-feature objects requires highly precise grasp pose estimation to ensure reliable in-gripper alignment. Vision-based approaches (e.g., YOLO-derived detectors) can localize the object but often fail to recover the exact grasp position under occlusion or partial views, while tactile-only methods lack global context; moreover, existing work rarely evaluates, in a task-specific and systematic manner, which fusion configurations are most suitable for this setting. Motivated by this gap, this paper investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper; initial grasp detection is outside the scope of this work. A deep-learning regression model based on YOLOv11 as a feature extractor was developed to estimate the grasp position using inputs from an RGB eye-in-hand camera and a tactile sensor. Within the defined experimental setup, the early RGB–tactile fusion model selected via NAS achieved the lowest summed lateral error of 1.02 mm, compared to tactile-only (1.35 mm) and RGB-only (1.31 mm) models trained under identical protocols. These results reflect performance under the dataset’s image-level partitioning and indicate improved lateral localization within the evaluated setup. The study does not claim a universally optimal fusion strategy, but provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.

Adam Michael Altenbuchner, Bsher Karbouj, Fabian Dilly et al. · 0 citations
Open access Aug 2026

A Unified Multi-Task Deep Learning Framework for Robotic Bin-Picking of Planar Objects

Automating random bin-picking in industrial robotics, where robots handle diverse and cluttered objects, remains challenging due to the complexity of object detection and pose estimation. While many solutions focus on free-form objects, systems specifically designed for planar objects are lacking. Planar objects pose unique challenges, as the commonly used point pair feature approach for free-form objects is ineffective due to their lack of distinctive geometric features. In this study, the proposed framework was implemented and evaluated using USB packs as a representative planar object case study. An innovative approach is introduced for the random bin-picking of planar objects by developing a multi-task model for instance segmentation and keypoint detection in 2D images. Geometric approach is then employed to estimate the 6D object pose for robotic grasping. Furthermore, a grasp candidate selection strategy is proposed to enable reliable grasping in cluttered industrial environments. Experimental results show that the proposed method achieved mAP50 values of 0.954, 0.800, and 0.926 for bounding box detection, instance segmentation, and keypoint detection, respectively, with a processing time of 2.7 ms. Future work will focus on integrating the framework into a digital twin system to support real-time monitoring, simulation, and optimization of automated manufacturing processes.

Ho Chi Minh, The-Thinh Pham, Tuan-Khanh Nguyen et al. · 0 citations