Skip to content
Conference

FG-Net: frequency-domain and Gaussian joint enhancement for 3D human pose estimation

Jul 2026 · International Conference on Image Processing and Intelligent Control · Vol 14262, pp. 142620Q - 142620Q-6 · 0 citations · 8 references
Engineering

Abstract

Transformers have achieved remarkable performance in video-based 3D human pose estimation, yet their high computational cost hinders deployment on resource-constrained devices. To balance accuracy and efficiency, this paper proposes an efficient plug-and-play joint enhancement framework, FG-Net, which integrates the Frequency Enhancement Module (FEM) and Gaussian Enhancement Module (GEM) to boost the performance of video pose Transformers. On this basis, FEM calibrates the semantic consistency of multi-scale features through frequency-domain detail enhancement and deformable spatial alignment, compensating for information loss caused by sampling. GEM constructs graph attention based on human skeletal topology, combined with temporal Gaussian smoothing and residual fusion, to adaptively strengthen joint structures, suppress noise and temporal jitter. The two modules work in synergy, enabling the model to achieve efficient inference while being more robust to low-quality video inputs such as occlusion and motion blur. Experimental results on the public dataset Human3.6M demonstrate that the proposed method achieves 39.87 mm MPJPE, obtaining state-of-the-art accuracy with lower computational complexity. The framework is generic and can be seamlessly integrated into mainstream video pose Transformers, providing an effective solution for real-time 3D human pose estimation in resource-constrained scenarios.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.