Aug 2026· Discover Computing· Vol 29· 0 citations· 21 references
TL;DR
A real-time point cloud processing and workpiece localization system that integrates multi-module optimization with a dynamic adaptive framework that sustains reliable performance under Gaussian noise up to 0.6 mm and occlusion levels up to 40%, confirming its viability for high-throughput industrial applications.
Abstract
The transition from automated to intelligent manufacturing increasingly relies on three-dimensional (3D) machine vision for robot guidance. Nevertheless, existing 3D vision systems still suffer from long processing delays, poor adaptability to changing environments, and inadequate pose accuracy in real industrial settings. To address these issues, this paper proposes a real-time point cloud processing and workpiece localization system that integrates multi-module optimization with a dynamic adaptive framework. The system acquires data through binocular stereo vision and incorporates optimized spatial filtering, PCA-based dimensionality reduction, and a machine learning-enhanced FAST feature detector. A hand-eye calibration model that includes distortion compensation achieves sub-millimeter mapping from image coordinates to the robot workspace. Experimental results on the public LineMOD dataset and a custom industrial bin-picking dataset show that the proposed system attains a 99% grasping success rate under favorable lighting and 96% under challenging conditions, with an average cycle time of 520 ms. Translation error reaches 0.38 mm under good lighting (rotation error: 0.61°, ADD: 0.47 mm).Comparative evaluations confirm substantial gains in speed, accuracy, and environmental robustness relative to existing commercial and academic systems. Furthermore, the system sustains reliable performance under Gaussian noise up to 0.6 mm and occlusion levels up to 40%, confirming its viability for high-throughput industrial applications.
The process of creating an integrated system for high-precision autonomous object grasping by an ABB IRB 140 industrial robot based on RGB-D perception and deep learning methods is presented. A fully functional real-time system for the resource-limited NVIDIA Jetson Nano platform is proposed, combining multi-angle 3D reconstruction of the scene with the exclusion of the central zone of the increased error, adaptive processing of depth data and neural network object detection based on YOLOv3. An algorithm for merging point clouds with weighted averaging in voxel representation and adaptive filtering by manipulator configuration has been developed; an ablation study of the contribution of the central zone mask, multi-angle averaging and adaptive filtering to the final capture accuracy has been conducted. The system achieves a gripping accuracy of 97.96 % with an average cycle time of 4.2 ± 0.8 s for orderly stacking and 94.74 % at 6.8 ± 1.5 s for disordered bales with an average localization error of 2.3 ± 1.5 cm. The reliability of the results is confirmed by statistical analysis and comparison with the methods of 6IMPOSE, PVN3D+ and GG-CNN. The advantage of the proposed approach in terms of capture accuracy and occlusion resistance at a comparable processing time is shown. The results obtained demonstrate the practical applicability of the system for autonomous manipulation in production environments with partial occlusion and variable scene geometry.
V. Meshcheryakov, S. Kondratyev, M. Kazakov· Journal of Instrument Engine...· 0 citations
Real-time and precise recognition of moving targets by industrial robots based on optical vision in dynamic production environments is key to realizing intelligent grasping and assembly. Existing detection methods still suffer from insufficient feature representation capability and difficulty in balancing detection speed and accuracy when dealing with rapid changes in target scale, motion blur, optical reflection interference and complex background noise. To address these issues, this paper proposes an improved YOLOv13 (DCA-YOLOv13) recognition model. The model introduces a deformable convolution module into the backbone network to enhance the geometric adaptation capability for non-rigidly deformed targets, and designs a lightweight channel attention mechanism to suppress complex background noise. At the same time, a multi-scale feature gold tower is integrated to improve sensitivity to small targets. Experiments conducted on a self-built optical dynamic workpiece image dataset show that the improved model achieves an enhanced mean average precision (mAP) of 94.6% and an inference speed of 82 FPS, meeting the real-time requirements of industrial sites. This model provides an efficient and stable optical visual perception scheme for robotic dynamic grasping.
Xi-Tao Song· European Conference on Elect...· 0 citations
The garment sewing industry faces persistent operator shortages and inconsistent seam quality. This work presents a vision-based robotic system for automating the perimeter-stitch, i.e. the Yun operation in shirt manufacturing, covering cuff, collar, and pocket-flap assembly. An improved Holistically-nested Edge Detection (HED) network is proposed, incorporating a directionally-aware Frequency-Spatial (FS) attention module that preserves spatial coordinate information and enhances boundary localisation in texture-rich garment images. A domain-specific dataset was constructed using garment-panel images captured under front-lit and backlit illumination conditions. The improved model achieves an Optimal Dataset Scale (ODS) F-measure of 0.8517, improving over the baseline HED model by 4.02 percentage points in ODS and 4.21 percentage points in Optimal Image Scale (OIS). Comparisons with representative edge-detection models, including PiDiNet, DexiNed, and EDTER, further demonstrate that the proposed model achieves a favourable balance between accuracy and efficiency. Stitch coordinates are generated in approximately 10 s via contour extraction and polygon approximation. The MS6MT six-degree-of-freedom (6-DOF) manipulator executes Cartesian-space linear and circular-arc trajectories derived from these coordinates through camera and eye-to-hand calibration. Physical experiments demonstrate stitch placement accuracy within ± 1 mm, validating the feasibility of the system under controlled experimental conditions.
Neng-Sheng Bao, Kewei Wang, Alessandro Simeone et al.· Journal of Intelligent Manuf...· 0 citations
The advancement of robotics has expanded applications across various sectors, increasing the need for reliable mapping and navigation in unfamiliar environments. Simultaneous Localization and Mapping (SLAM) enables mobile robots to estimate their position while constructing an environmental map, while RGB-D SLAM combines visual and depth information for three-dimensional perception. This study implements and evaluates an RGB-D SLAM system using a Microsoft Kinect for Xbox 360 integrated with the Robot Operating System (ROS) on a differential-drive mobile robot. The system was evaluated through Gazebo simulations and real-world experiments under four conditions: tidy indoor, cluttered indoor, dynamic indoor, and outdoor environments. The system successfully generated 2-D octomaps and 3-D meshes across the tested conditions. Quantitative evaluation showed that the selected RTAB-Map configuration, with a queue size of 20 and an odometry maximum rate of 10, resulted in mean memory-update and data-compression times of 56.9 s and 10.6 s, respectively. The Kinect exhibited an effective mapping range of approximately 1–3 m, with a blind region below 1 m. The results demonstrate the feasibility of combining a low-cost RGB-D sensor with ROS-based SLAM for practical 2-D and 3-D mapping while also highlighting limitations associated with occlusion, dynamic objects, and sensor range.
Fahmizal, Priyova Muhammad Rafief, Rico Agustiawan et al.· Applied Sciences· 0 citations
This research presents a low-cost automated nut sorting system developed through laboratory testing to provide an affordable automation solution for small and medium enterprises (SMEs). The system integrates a fixed-vision camera with a Dobot Magician robotic arm, utilizing Python, OpenCV, and homography-based coordinate transformation for precise positioning. Performance was evaluated under three controlled lighting conditions with 25 samples each. Results indicate that lighting intensity significantly affects accuracy: low lighting (27.25 lux) yielded only 16% accuracy (mu=0.16, sigma approximately 0.3666), while high lighting (481.78 lux) suffered from overexposure and reflections. In contrast, optimal laboratory conditions (129.71 lux) achieved 100% classification accuracy (mu=1, sigma=0), demonstrating perfect consistency and stability. The study concludes that while the system offers a high-efficiency, budget-friendly alternative for SMEs, maintaining controlled, optimal illumination is critical for operational success. These findings provide a technical foundation for implementing cost-effective robotic sorting in real-world SME environments where high-cost sensor arrays are not feasible.
Phisit Srinoi, Surasit Phokha, Viroch Sukontanakarn et al.· Bulletin of Electrical Engin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.