Skip to content

An improved, scalable and modular framework for vision-based robotic grasp detection

Jul 2026 · International Journal of Machine Learning and Cybernetics · Vol 17 · 0 citations · 28 references
Computer Science

TL;DR

An EfficientNet based scalable and modular model has been presented for the robotic grasp detection task, which builds upon the EfficientPose model by proposing subnets for the robotic grasp detection task.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

A reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN), is proposed, offering a scalable and adaptable solution for contact-rich manipulation tasks.

Amir Arsalan Nematollahi, Shayan Ahmadi, M. T. Masouleh et al. · 0 citations
Conference Open access Aug 2026

VLEG: Embodied Vision-Language Grasping for a Quadruped Manipulator

Open-vocabulary grasping on a quadruped manipulator requires more than recognizing the target object. The robot must also select a grasp pose that is both consistent with the task semantics and reliable to execute under body motion and viewpoint changes. In this paper, we present VLEG, an embodied vision-language grasping framework for quadruped manipulators that explicitly incorporates body motion into grasp decision making. Our method guides the robot to continuously adjust its body pose during approach and optimize local observations before grasping, thereby improving perception quality. For grasp decision making, instead of using a coarse single-stage filtering strategy, we design a multi-stage and multi-criteria grasp selection mechanism based on geometric grasp candidates. This mechanism jointly considers physical feasibility and task consistency. We implement the complete system on an onboard Jetson platform and conduct extensive real-world experiments on a quadruped robot equipped with a manipulator, covering tabletop, low-platform, ground-level, and outdoor raised-platform scenes. The results validate the deployability of VLEG in real-world quadruped manipulation scenarios, as well as its robust grasping ability and task-aware decision-making capability across the tested object categories.

Yu-Xing Ji, Fei Meng, Zishang Ji et al. · 0 citations
Open access Aug 2026

Kitchen robotic manipulation utilizing foundation models

A modular perception pipeline for household manipulation tasks, with a focus on dishware handling in kitchen environments, that integrates open-vocabulary object detection, multi-view segmentation, instance-aware 3D reconstruction, and a 2D-3D feature fusion strategy for 6D pose estimation and grasp planning is presented.

Myung-Hwan Jeon, Sankalp Yamsani, Joohyung Kim · 0 citations
Aug 2026

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.

Hui Zhang, Yue Wang, Kang An et al. · 0 citations
Open access Aug 2026

A Unified Multi-Task Deep Learning Framework for Robotic Bin-Picking of Planar Objects

An innovative approach is introduced for the random bin-picking of planar objects by developing a multi-task model for instance segmentation and keypoint detection in 2D images and a grasp candidate selection strategy is proposed to enable reliable grasping in cluttered industrial environments.

The-Thinh Pham, Tuan-Khanh Nguyen, Chi-Cuong Tran et al. · 0 citations
Jul 2026

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

This work proposes SeededGrasp, a novel data-efficient framework that enables a VLM to predict a seed point to be used as conditioning for a subsequent lightweight grasp-generation model, enabling multi-embodiment support while bypassing the need for expensive end-to-end training.

Yang Xu, Gurpreet Singh Mukker, Raymond Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.