Skip to content
Open access

WaveletkAN Pose Refinement for Robot Bin Picking Driven by Cluttered RGB-D Point Clouds

Jul 2026 · Journal of Applied Automation Technologies · Vol 4 · 0 citations · 27 references

TL;DR

WaveletkAN is a pose-refinement model based on RGB-D point clouds proposed in this paper to correct object pose hypotheses for dense bins with occlusion, specular depth noise, and self-similar industrial parts and can achieve robust and deployable pose correction for RGB-D robotic bin-picking systems.

Abstract

Grasping cluttered boxes is still limited by the gap between rough 6-degree-of-freedom perception and the millimeter-level alignment required for robot execution. WaveletkAN is a pose-refinement model based on RGB-D point clouds proposed in this paper to correct object pose hypotheses for dense bins with occlusion, specular depth noise, and self-similar industrial parts. The method represents each candidate as a local observed point set, a CAD-derived canonical set, and a compact bin-context tensor, then predicts a residual rigid motion via Kolmogorov-Arnold network layers parametrized by learnable wavelet atoms. A confidence-weighted correspondence field and a residual SE (3) update can be used with the traditional proposal generator and grasp planner. WaveletkAN reduced the median translation error from 7.8 mm to 2.9 mm and the median rotation error from 5.6 degrees to 1.9 degrees after refinement in a simulated-real mixed benchmark with 18 object categories and 42,600 evaluated hypotheses. The successful pick rate in dense clutter rose to 93.1 %, and the mean refinement latency of the industrial GPU was still 11.6 ms. Ablation experiments showed that removing the wavelet basis increased ADD-S by 31.7%, and omitting context gating reduced top-1 executable pose recall by 6.8 %. Based on the above experiments, multi-scale functional parameterization can achieve robust and deployable pose correction for RGB-D robotic bin-picking systems.

Read PDF

Similar papers

Open access Aug 2026

Robot-Centric Elevation Map Completion with Sensor Geometry-Aware Augmentation and Uncertainty Estimation

Robot-centric elevation maps built from onboard sensing are always incomplete: occlusions, a limited field of view, and range limits leave large unobserved regions that traversability analysis and motion planning must still reason about. We present a supervised framework that completes these maps and reports a per-cell uncertainty. Its core is a ray-cone augmentation that removes angular sectors anchored at the sensor origin during training; unlike the random masks of image inpainting, these sectors match the coverage gaps of real deployments, such as camera failures or reduced camera configurations. Partial maps generated from four depth cameras along legged-robot trajectories in the TartanGround dataset are paired with dense ground truth, yielding 32,329 samples across five outdoor environments. An encoder–decoder network is trained with a masked β-NLL loss and evaluated with a five-fold leave-one-environment-out protocol. The augmentation lowers the completion error on missing sensor sectors by 8.3 to 9.7%, depending on the sector width, at no measurable cost on uncorrupted partial inputs. The completed maps reach a hole root-mean-square error of 2.86 m, a 45% improvement over the strongest classical interpolation baseline.

J. Goga, Michal Kovac, Martin Dekan et al. · 0 citations
Open access Aug 2026

Vision-Guided Robotic Bin-Picking of Disordered Workpieces via Image-Matching Pose Estimation

Robotic bin-picking of disordered, randomly stacked workpieces remains challenging because reliable grasping depends on an accurate estimate of object pose, yet many established solutions require high-precision 3D sensing, detailed object models, or large annotated datasets that raise the cost and effort of deployment on a new production line. This work presents a complete binocular vision framework that estimates workpiece pose by image matching and executes vision-guided grasping on a 6-DOF manipulator. A pose-annotated multi-view template library is constructed automatically through robot-driven image acquisition and compressed by a coarse-to-fine clustering scheme, and object pose is estimated by discriminative template matching with rigid refinement. To characterize the geometric reliability of the matched poses, an offline cross-modal analysis relates the 2D templates to a 3D reference model of the object and measures their agreement through region and contour reprojection metrics. Grasp configurations are then generated under orientation and collision constraints and corrected online by closed-loop visual feedback. Experiments on two representative workpieces show template-matching accuracy of 89–90% against classical and learned similarity measures, and grasp success between 81 and 87% across single-object and mixed scenes, outperforming the GraspNet baseline under the tested conditions. The framework offers an accurate and deployment-friendly route to robotic bin-picking.

Abdulrahman Usman Wunti, Ling-Xin Yu, Guangwei Li et al. · 0 citations
Conference Aug 2026

Camera-Based Synthetic Completion of Point Clouds Captured with Quadruped Robot

LiDAR-based perception systems commonly used in mobile robots often struggle to accurately capture transparent or reflective surfaces such as glass, leading to incomplete point clouds and degraded understanding of the scene. This paper presents a camera-based synthetic point cloud completion method designed to address these limitations. The proposed approach integrates RGB images with LiDAR measurements using a quadruped robot equipped with synchronized sensors. A deep learning model based on YOLOv26 is trained to detect and segment window regions in camera images. The resulting semantic information is projected onto corresponding LiDAR data to identify areas with missing geometry. For each detected region, planar surfaces are estimated and synthetic points are generated within these boundaries to reconstruct the missing structures. Experimental evaluation conducted on a real-world dataset demonstrates that the method significantly improves the completeness and consistency of point clouds, particularly in areas containing glass surfaces. The enhanced maps provide more accurate 3D representations, which can improve navigation, obstacle avoidance, and path planning in robotics.

J. Koszyk, Bartosz Hyla, Ł. Ambroziński · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations
Open access Sep 2026

MRPose: Multi‐Robot Relative 6D Pose Estimation for Unseen Objects From RGB Images

Relative pose estimation under the query–reference paradigm has emerged as a practical alternative for estimating the 6D pose of unseen objects without relying on CAD models or extensive annotations. However, many existing approaches remain difficult to deploy in practice, as their reliance on geometric matching makes robustness highly sensitive to matching quality. In this paper, we present MRPose, a framework for relative 6D object pose estimation under a multi‐robot setting, where query and reference views are independently captured by different robots. Without requiring CAD models, shared calibration across robots or pose annotations, MRPose directly estimates the full relative transformation between unseen objects from RGB images. Built upon the power of vision foundation models, MRPose introduces a multiplex feature learning strategy that jointly exploits patch‐level appearance correspondence and sparse geometric cues through attention‐based cross‐view feature aggregation. Instead of relying on traditional geometric solvers such as PnP or essential matrix estimation, the proposed framework directly regresses the full relative 6D pose in an end‐to‐end manner. Extensive experiments on the LINEMOD and YCB‐Video benchmarks demonstrate that MRPose compares favourably with existing RGB‐based relative pose estimation methods. Real‐world experiments further validate its applicability in practical scenarios.

Jia-Le Ren, Ming-Xing Tan, Meng-Yuan Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.