Skip to content
Open access

FreeTrack6d: a training-free framework for 6D pose tracking and robotic grasping of moving objects on conveyor belts

Aug 2026 · Engineering Research Express · Vol 8, pp. 165233 · 0 citations · 33 references
Physics

TL;DR

FreeTrack6D, a training-free unified segmentation and 6D pose tracking framework integrating frame-wise mask generation, multi-hypothesis pose refinement, motion prediction, observation gating, and visual-servo-based robot control, forms a closed-loop training-free dynamic grasping framework.

Abstract

For robotic dynamic grasping of moving objects in conveyor-belt scenarios, accurate and robust 6D pose estimation and tracking are essential for reliable grasping. However, existing deep-learning-based methods usually rely on large amounts of supervised data for specific objects or categories, which limits their generalization, deployment efficiency, and flexibility for rapid object changeover in industrial applications. To address these challenges, this paper proposes FreeTrack6D, a training-free unified segmentation and 6D pose tracking framework. Relying only on the CAD model of the target object, FreeTrack6D can be applied to dynamic tracking and grasping of unseen objects without additional object-specific training. Specifically, an adaptive multi-cue mask generation module is first introduced to generate frame-wise target masks in real time, which provides target-region constraints for initial pose registration and subsequent pose refinement. This helps reduce the influence of background interference and target-region misalignment caused by rapid motion. Based on the generated mask, RGB-D observations, and the CAD model, FoundationPose is used for initial 6D pose registration and subsequent pose refinement. To improve tracking robustness under large inter-frame motion and rotational variations, a Kalman-guided multi-hypothesis refinement strategy is further designed, where multiple candidate poses predicted from historical motion states are refined and selected according to mask consistency. In addition, a Pose Consistency-aware Association and Gating mechanism is developed to reject abnormal detections, protect the filter state, and trigger re-initialization when consecutive mismatches occur. By integrating frame-wise mask generation, multi-hypothesis pose refinement, motion prediction, observation gating, and visual-servo-based robot control, FreeTrack6D forms a closed-loop training-free dynamic grasping framework.

Read PDF

Similar papers

Jul 2026

RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes

Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computationally expensive, while current trackers struggle with fast motion and complete occlusion due to their reliance on continuous visibility. To address these challenges, we present RRTrack, an efficient, recoverable object 6D pose tracker that enables robust tracking through fast motion and target disappearance--reappearance. RRTrack introduces a 2D--6D closed-loop tracking strategy that integrates memory-based video object segmentation (VOS) with 6D pose refinement. The 2D branch maintains target localization, and the 6D branch verifies geometric consistency before memory updates. In addition, a DINOv2-based dual-bank template matching module is developed to recover lost targets by jointly exploiting offline synthetic templates and online observation anchors while maintaining real-time efficiency. We also introduce a synthetic RGB-D benchmark comprising three robotic scenarios with fast motion and full occlusion. Experimental results on the synthetic benchmark demonstrate that RRTrack improves equal-subset mean ADD-S AR by 66.3\% and ADD-S AUC by 65.7\% over FoundationPose while achieving 55.2 FPS. Real-world experiments further validate the robustness of RRTrack under noisy sensing conditions. Project page: https://github.com/7kevin24/RRTrack

Jun-Yue Li, Ye Zheng, Yifan Chen et al. · 0 citations
Open access Jul 2026

Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions

Real-time 6-DoF object pose tracking is essential for many robotics applications, and several approaches exist. Yet even today's approaches remain unreliable under temporary full occlusions and rapid object motions. Once tracking is lost, most methods struggle to detect the failure and recover automatically, often requiring manual re-initialization. In this paper, we address the problem of robust model-based 6-DoF tracking of unseen objects from RGB-D data, especially in scenarios with occlusion and fast motion. We propose a novel method that combines efficient learning-based keypoint matching with optimization-based alignment and introduces a novel failure detection and recovery module. Our system monitors pose reliability, detects tracking divergence or occlusions, and performs a global re-detection and pose estimation step that robustly verifies recovery candidates before resuming tracking. Our evaluation on standard tracking benchmarks and on a new dataset of occluded and fast-moving scenes shows that our method matches state-of-the-art accuracy on easy tracking sequences, maintains high tracking speed at 57.6 frames per second, and provides the most robust tracking performance under challenging conditions. Thus, we believe that our approach is a relevant step forward in robust 6-DoF object tracking from RGB-D data.

Balázs Opra, Léo Ghafari, Thomas Stewart et al. · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations
Open access 2026

Grasp Pose Estimation of Articulated Objects Based on Semantic and Geometric Feature Fusion

A deep learning-based grasp estimation model designed to enable robotic manipulation with articulated objects that incorporates the attention-based semantic and geometric feature fusion (ASGF) module improved the grasp success rate in the evaluated setting.

Dongwoo Lee, Yeongmin Kim, Seong Bin Jo et al. · 0 citations
Open access Aug 2026

A Unified Multi-Task Deep Learning Framework for Robotic Bin-Picking of Planar Objects

An innovative approach is introduced for the random bin-picking of planar objects by developing a multi-task model for instance segmentation and keypoint detection in 2D images and a grasp candidate selection strategy is proposed to enable reliable grasping in cluttered industrial environments.

The-Thinh Pham, Tuan-Khanh Nguyen, Chi-Cuong Tran et al. · 0 citations
Aug 2026

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.

Hui Zhang, Yue Wang, Kang An et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.