A Kinematics-Aware Decoupled Learning Network for robust 3D hand pose estimation and a task-decoupled hierarchical multitask framework, which separates 2D joint localization from depth estimation and incorporates a dedicated multitask learning strategy for depth regression.
Abstract
Due to the complexity of hand kinematics and self-occlusion, existing 3D hand pose estimation methods based on single depth images struggle to comprehensively model the topological dependencies among hand joints. Furthermore, traditional hierarchical multitask architectures enforce a shared feature space for both 2D joint localization and depth estimation, which can induce mutual interference. To address these challenges, we propose a Kinematics-Aware Decoupled Learning Network (KAD-Net) for robust 3D hand pose estimation. Specifically, we first design a Finger Topology Constraint (FTC) module to enhance the representation of distal joints. This module utilizes three consecutive finger joints to construct a local kinematic representation to impose topological constraints, which supplements the kinematic features of the distal joints. The FTC module leverages the structural context from visible joints to assist in locating occluded distal joints, thereby improving robustness to occlusion. Additionally, we propose a task-decoupled hierarchical multitask framework. This framework separates 2D joint localization from depth estimation and incorporates a dedicated multitask learning strategy for depth regression, effectively isolating the UV and depth features to mitigate mutual interference and negative transfer. Extensive experiments demonstrate that KAD-Net outperforms existing methods on several benchmark datasets (ICVL, NYU, and MSRA), achieving state-of-the-art accuracy in 3D hand pose estimation. Potential applications of KAD-Net include human-computer interaction, virtual reality and gesture-based control systems.
Existing head pose estimation (HPE) methods focus on utilizing the information of one single image and ignore cross-viewpoint pose consistency. To remedy this, we propose a dual-branch self-correcting pose consistency network (DSPC-Net) for HPE. The idea is to construct flipped image pairs as intrinsic pose constraints...
Chun Liu, Tie-Cheng Song, Feng Yang et al.· IEEE Signal Processing Lette...· 0 citations
Monocular 6D pose estimation remains challenging due to low-quality 3D coordinate maps caused by occlusion, texture-less surfaces, and spatial detail loss in encoder-decoder networks. This paper presents an efficient monocular framework that improves pose accuracy by enhancing the quality of dense 3D coordinate maps vi...
Ming-Rui Luo, Ming-Hao Chen, Xue-Wei Cao et al.· 2026 IEEE 22nd International...· 0 citations
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object...
Hong-Li Xu, Zhao-Wei Lu, Jun-Wen Huang et al.· 0 citations
This work presents FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoders supervised on object-scene overlap and reducing supervision requirements and memory overhead.
Amir Hamza, Davide Boscaini, Fabio Poiesi· 0 citations
Transformers have become dominant in 3D Human Pose Estimation (HPE). However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work we leverage recent advances in state space models and utilize Mamba for high-qua...
Xin-Yi Zhang, Qi-Qi Bao, Wen-Ming Yang et al.· ACM Transactions on Multimed...· 0 citations
Estimating the 6D pose of unseen objects from depth observations is fundamental for robotic perception, particularly for texture-less objects where RGB appearance provides limited cues. Despite rich geometric information provided by depth images, depth-only model-free 6D pose estimation remains largely underexplored. I...
Qing-Yang Zhou, Zi-Heng Li, Qing-Zhe Li et al.· IEEE Robotics and Automation...· 0 citations
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.