Jul 2026· 2026 6th International Conference on Intelligent Communications and Computing (ICICC)· pp. 361-365· 0 citations· 23 references
Abstract
While most robotic research focuses on household tasks such as bus table arrangement and cloth folding, numerous manipulation tasks remain challenging for industrial applications, particularly the grasping and transportation of scattered workpieces. In this paper, we propose SGDIFF, a vision-guided grasping architecture that integrates a pre-trained vision-language model (VLM) with flow matching-based diffusion model. Given an RGB-D image of a tabletop scene, the VLM first detects each workpiece, outputs its bounding box, and assigns a unique ID. A point cloud is then generated for each detected instance. Subsequently, flow matching model iteratively refines the initially noisy gripper pose to a stable and collision-free grasp for each target workpiece. The proposed method eliminates the need for object-specific models and enables efficient multi-object grasping in cluttered industrial environments. The experiments validates the effectiveness of combining semantic understanding with precise pose refinement for robust industrial automation.
Robotic bin-picking of disordered, randomly stacked workpieces remains challenging because reliable grasping depends on an accurate estimate of object pose, yet many established solutions require high-precision 3D sensing, detailed object models, or large annotated datasets that raise the cost and effort of deployment on a new production line. This work presents a complete binocular vision framework that estimates workpiece pose by image matching and executes vision-guided grasping on a 6-DOF manipulator. A pose-annotated multi-view template library is constructed automatically through robot-driven image acquisition and compressed by a coarse-to-fine clustering scheme, and object pose is estimated by discriminative template matching with rigid refinement. To characterize the geometric reliability of the matched poses, an offline cross-modal analysis relates the 2D templates to a 3D reference model of the object and measures their agreement through region and contour reprojection metrics. Grasp configurations are then generated under orientation and collision constraints and corrected online by closed-loop visual feedback. Experiments on two representative workpieces show template-matching accuracy of 89–90% against classical and learned similarity measures, and grasp success between 81 and 87% across single-object and mixed scenes, outperforming the GraspNet baseline under the tested conditions. The framework offers an accurate and deployment-friendly route to robotic bin-picking.
Abdulrahman Usman Wunti, Ling-Xin Yu, Guangwei Li et al.· Applied Sciences· 0 citations
Industrial robots in small-batch manufacturing and human–robot collaborative workstations are increasingly required to perform grasping and sorting tasks on workpieces with natural language instructions. Unlike generic object grasping, industrial workpieces may contain local surface defects such as dents, scratches, and geometric irregularities, which affect both target selection and grasp stability. This paper presents a geometric point-cloud perception and defect-aware grasping framework that leverages 3D geometric cues measured from point clouds to detect surface defects and guide robotic manipulation. Rather than learning an end-to-end visuomotor policy, the proposed framework introduces an explicit geometric reasoning layer between language-level task specification and robotic grasp execution. The language model is used only to convert instructions into structured task attributes, whereas defect perception, target ranking, and grasp-contact evaluation are performed using measurable 3D geometric evidence. These components require no task-specific training on annotated industrial defect or grasping datasets, making the formulation potentially useful for small-batch manufacturing scenarios in which large annotated defect datasets are unavailable. A shared defect-response representation supports both instance-level target selection and contact-region screening. We evaluate the proposed method on a controlled prototype tabletop setup, achieving an average defect instance-level F1-score of 0.89, a target grounding accuracy of 0.89, and a grasp success rate of 84.0%.
Yufeng Li, Hai-Feng Yu, Xin Su et al.· Mathematics· 0 citations
An innovative approach is introduced for the random bin-picking of planar objects by developing a multi-task model for instance segmentation and keypoint detection in 2D images and a grasp candidate selection strategy is proposed to enable reliable grasping in cluttered industrial environments.
The-Thinh Pham, Tuan-Khanh Nguyen, Chi-Cuong Tran et al.· Journal of Technical Educati...· 0 citations
A modular perception pipeline for household manipulation tasks, with a focus on dishware handling in kitchen environments, that integrates open-vocabulary object detection, multi-view segmentation, instance-aware 3D reconstruction, and a 2D-3D feature fusion strategy for 6D pose estimation and grasp planning is presented.
A deep learning-based grasp estimation model designed to enable robotic manipulation with articulated objects that incorporates the attention-based semantic and geometric feature fusion (ASGF) module improved the grasp success rate in the evaluated setting.
Dongwoo Lee, Yeongmin Kim, Seong Bin Jo et al.· IEEE Access· 0 citations
A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.
Hui Zhang, Yue Wang, Kang An et al.· Signal, Image and Video Proc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.