Back to feed
Open access

Automated Model Selection for Task-Specific RGB–Tactile Fusion in In-Hand Grasp Pose Estimation

2026 · IEEE Access · Vol 14, pp. 125611-125617 · 0 citations · 18 references

Abstract

Robotic grasping of small, low-feature objects requires highly precise grasp pose estimation to ensure reliable in-gripper alignment. Vision-based approaches (e.g., YOLO-derived detectors) can localize the object but often fail to recover the exact grasp position under occlusion or partial views, while tactile-only methods lack global context; moreover, existing work rarely evaluates, in a task-specific and systematic manner, which fusion configurations are most suitable for this setting. Motivated by this gap, this paper investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper; initial grasp detection is outside the scope of this work. A deep-learning regression model based on YOLOv11 as a feature extractor was developed to estimate the grasp position using inputs from an RGB eye-in-hand camera and a tactile sensor. Within the defined experimental setup, the early RGB–tactile fusion model selected via NAS achieved the lowest summed lateral error of 1.02 mm, compared to tactile-only (1.35 mm) and RGB-only (1.31 mm) models trained under identical protocols. These results reflect performance under the dataset’s image-level partitioning and indicate improved lateral localization within the evaluated setup. The study does not claim a universally optimal fusion strategy, but provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.

Read PDF