Skip to content

KAD-Net: Kinematics-Aware Decoupled Learning for Robust 3D Hand Pose Estimation from a Single Depth Image

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

A Kinematics-Aware Decoupled Learning Network for robust 3D hand pose estimation and a task-decoupled hierarchical multitask framework, which separates 2D joint localization from depth estimation and incorporates a dedicated multitask learning strategy for depth regression.

Abstract

Due to the complexity of hand kinematics and self-occlusion, existing 3D hand pose estimation methods based on single depth images struggle to comprehensively model the topological dependencies among hand joints. Furthermore, traditional hierarchical multitask architectures enforce a shared feature space for both 2D joint localization and depth estimation, which can induce mutual interference. To address these challenges, we propose a Kinematics-Aware Decoupled Learning Network (KAD-Net) for robust 3D hand pose estimation. Specifically, we first design a Finger Topology Constraint (FTC) module to enhance the representation of distal joints. This module utilizes three consecutive finger joints to construct a local kinematic representation to impose topological constraints, which supplements the kinematic features of the distal joints. The FTC module leverages the structural context from visible joints to assist in locating occluded distal joints, thereby improving robustness to occlusion. Additionally, we propose a task-decoupled hierarchical multitask framework. This framework separates 2D joint localization from depth estimation and incorporates a dedicated multitask learning strategy for depth regression, effectively isolating the UV and depth features to mitigate mutual interference and negative transfer. Extensive experiments demonstrate that KAD-Net outperforms existing methods on several benchmark datasets (ICVL, NYU, and MSRA), achieving state-of-the-art accuracy in 3D hand pose estimation. Potential applications of KAD-Net include human-computer interaction, virtual reality and gesture-based control systems.

View source

Similar papers

2026

DSPC-Net: A Dual-Branch Self-Correcting Pose Consistency Network for Head Pose Estimation

Existing head pose estimation (HPE) methods focus on utilizing the information of one single image and ignore cross-viewpoint pose consistency. To remedy this, we propose a dual-branch self-correcting pose consistency network (DSPC-Net) for HPE. The idea is to construct flipped image pairs as intrinsic pose constraints...

Chun Liu, Tie-Cheng Song, Feng Yang et al. · 0 citations
Conference Aug 2026

Advanced Monocular 6D Pose Estimation with Enhanced 3D Coordinate Maps via Feature Fusion and Deformable Decoupled Regression

Monocular 6D pose estimation remains challenging due to low-quality 3D coordinate maps caused by occlusion, texture-less surfaces, and spatial detail loss in encoder-decoder networks. This paper presents an efficient monocular framework that improves pose accuracy by enhancing the quality of dense 3D coordinate maps vi...

Ming-Rui Luo, Ming-Hao Chen, Xue-Wei Cao et al. · 0 citations
Preprint Sep 2026

LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation

Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object...

Hong-Li Xu, Zhao-Wei Lu, Jun-Wen Huang et al. · 0 citations
Preprint Aug 2026

Foundational feature fusion for conditional flow matching in 6D pose estimation

This work presents FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoders supervised on object-scene overlap and reducing supervision requirements and memory overhead.

Amir Hamza, Davide Boscaini, Fabio Poiesi · 0 citations
Open access Sep 2026

Pose Magic++: Integrating Mamba and HyperGCN for Efficient and Temporally Consistent 3D Human Pose Estimation

Transformers have become dominant in 3D Human Pose Estimation (HPE). However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work we leverage recent advances in state space models and utilize Mamba for high-qua...

Xin-Yi Zhang, Qi-Qi Bao, Wen-Ming Yang et al. · 0 citations
Nov 2026

RGA6D: Regional Geometry-Aware Correspondence Reasoning for Depth-Only Model-Free 6D Pose Estimation

Estimating the 6D pose of unseen objects from depth observations is fundamental for robotic perception, particularly for texture-less objects where RGB appearance provides limited cues. Despite rich geometric information provided by depth images, depth-only model-free 6D pose estimation remains largely underexplored. I...

Qing-Yang Zhou, Zi-Heng Li, Qing-Zhe Li et al. · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.