2026· International Journal of Advanced Computer Science and Applications· Vol 17· 0 citations· 15 references
TL;DR
This study presents a deep learning approach for reconstructing the default canonical pose of non-rigid objects from single depth images that combines short-range and long-range feature extraction with the original depth input to capture both local geometric details and global structural information.
Abstract
Reconstructing the canonical pose of non-rigid objects from arbitrary depth observations is an important problem in robotic vision, particularly for systems that must perceive, track, and interact with deformable objects in dynamic environments. In robotics and SLAM-related perception, depth cameras are widely used to support object recognition, spatial understanding, scene mapping, and motion analysis. However, non-rigid deformation remains challenging because the same object may appear in significantly different poses, making reliable object-level representation and tracking difficult. In this study, we present a deep learning approach for reconstructing the default canonical pose of non-rigid objects from single depth images. The proposed model combines short-range and long-range feature extraction with the original depth input to capture both local geometric details and global structural information. By transforming arbitrary posed observations into a consistent canonical representation, the method supports more stable shape understanding, pose normalization, and object-level perception for robotic systems operating in real-world environments. This is particularly relevant to robotic vision tasks involving human motion analysis, deformable-object tracking, manipulation, and semantic mapping. The model is trained on synthetic human datasets and evaluated on synthetic human, real human, and animal datasets. Experimental results demonstrate improved retrieval accuracy compared with existing methods, showing that the proposed approach can generalize across different non-rigid categories and sensing conditions. These findings highlight the potential of canonical pose reconstruction as a useful component for intelligent robotic perception, depth-based scene interpretation, and SLAM-aware systems that require robust understanding of deformable objects.
6D object pose estimation is fundamental to robotic manipulation and automation. Recent zero-shot methods have significantly improved generalization to unseen objects, but most still rely on large-scale pretrained models with substantial GPU computation and memory demands. These requirements complicate deployment on ro...
Yi-Xuan Liang, William Chen, Yunan Wang et al.· 0 citations
State-of-the-art 3D reconstruction models, whether from visual, range, or both, tend to underperform on thin objects. This is partially due to the small amount of space such objects occupy in RGB images and in 3D point clouds. To test the extent of their errors, we collected the first thin object dataset comprising of...
Shania Guo, Yeongsik Seo, Andrew Fu et al.· 1 citation
Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle in such scenes: a missed object is never reconstructed, a merged detection can fuse two objects, and separately reconstru...
Dongwon Son, Junhyek Han, Yoon-Je Cho et al.· 0 citations
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations
LoFG is presented, a localization-oriented Feature Gaussian representation that unifies both stages within a single Gaussian scene and improves the robustness of sparse initialization and the accuracy of dense refinement, demonstrating potential for localization applications in AR, robotics, and visual navigation syste...
Zhen-Dong Xiao, Zi-Ling Wen, Jun Yin et al.· Multimedia Systems· 0 citations