Through-Wall Multiperson 3-D Pose Estimation With MIMO Radar in Large-Scale Scenarios
Abstract
The 3-D human pose estimation is a key task in the fields of computer vision and wireless sensing. Compared with optical sensors, through-wall (TW) radar can penetrate nonmetallic obstacles and capture reflected signals from targets, making it highly promising for applications in visually constrained environments. However, severe signal attenuation and low spatial resolution of TW radar make accurate human pose reconstruction still highly challenging. To this end, we propose an end-to-end multiperson 3-D pose estimation method based on multi-input multi-output (MIMO) TW radar and transformer (PERT). In PERT, we first extract fused multiscale features from horizontal and vertical radar heatmap sequences using a feature extractor. Subsequently, to enhance the focusing ability on effective regions in large-scale scenes and reduce computational overhead, we propose a signal-guided foreground selector (SFS) that leverages a signal-guided salience supervision mechanism to guide the selector in selecting tokens related to the targets. Finally, the spatiotemporal pose (STP) transformer module extracts fine-grained pose features from foreground tokens using an attention mechanism and predicts the positions of 3-D keypoints via a decoder. Experimental results demonstrate that PERT outperforms all baseline methods, achieving an average localization error of 6.20 cm in a large-scale $6\times 15$ m scene behind a 24-cm cement wall.