Skip to content
Open access

Pose Magic++: Integrating Mamba and HyperGCN for Efficient and Temporally Consistent 3D Human Pose Estimation

Sep 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 41 references

Abstract

Transformers have become dominant in 3D Human Pose Estimation (HPE). However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work we leverage recent advances in state space models and utilize Mamba for high-quality and efficient long-range modeling. Nonetheless, Mamba still faces challenges in precisely exploiting local dependencies between joints. To address these issues, we propose a novel attention-free integrated spatiotemporal architecture named Integrated Mamba-HyperGCN (Pose Magic++). This architecture introduces local enhancement with HyperGCN by capturing relationships between neighboring joints and their synergies, producing new representations to complement Mamba's outputs. By adaptively fusing representations from Mamba and HyperGCN, Pose Magic++ demonstrates superior capability in learning the underlying 3D structure. To meet the requirements of real-time inference, we also provide a fully causal version. Extensive experiments show that Pose Magic++ achieves new state-of-the-art (SOTA) results ( \(\downarrow 1.6mm\) ) while saving \(74.1\%\) FLOPs. Notably, it significantly improves the estimation accuracy of peripheral joints without incurring extra computational cost. In addition, Pose Magic++ exhibits optimal motion consistency and strong ability to generalize to unseen sequence lengths.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.