This work proposes Multi-Hypothesis Normalizing Flow Pose Generator (MH-NFPG), which models pose distributions from radar point clouds using a conditional normalizing flow that transforms a Laplace base distribution into an expressive posterior, generated in parallel through a single forward pass.
Abstract
Sparse and noisy millimeter-wave radar point cloud observations often correspond to multiple plausible human poses, making deterministic pose estimation fundamentally ill-posed. Yet existing radar methods remain deterministic, collapsing this ambiguity into a single estimate. Diffusion-based alternatives can model multi-hypothesis distributions but require costly sequential denoising for each distribution sample and lack calibrated uncertainty. We propose Multi-Hypothesis Normalizing Flow Pose Generator (MH-NFPG), which models pose distributions from radar point clouds using a conditional normalizing flow. Specifically, we combine a spatiotemporal transformer backbone with a normalizing flow that transforms a Laplace base distribution into an expressive posterior, generated in parallel through a single forward pass. Leveraging this efficiency, we outperform diffusion-based alternatives in calibration across three radar benchmarks (MM-Fi, mmRadPose, mRI), improve pose accuracy on two, and match it on the third, while achieving over 20x faster inference for applications and reducing calibration error by up to 85%. We find that calibration degrades substantially for diffusion models, whereas our flow-based approach maintains reliable coverage, also in cross-environment settings. These results demonstrate normalizing flows as a practical alternative to diffusion models for real-time, uncertainty-aware radar pose estimation. Our code will be made publicly available.
Dense metric depth prediction from cameras and millimeter-wave radar offers a cost-effective sensing solution for autonomous systems. However, radar measurements are inherently sparse and susceptible to clutter, multipath reflections, and projection errors. While aggregating multiple radar frames provides denser metric cues, it also introduces temporal misalignment and dynamic-object interference. Directly propagating such unreliable measurements can therefore corrupt large regions of the predicted depth map. To address this issue, we propose RbFT-Net, an end-to-end rectify-before-fuse framework for multi-frame 4D radar-camera depth completion. Rather than assuming accumulated radar returns to be accurate, RbFT-Net treats them as noisy temporal anchor candidates. An image-conditioned rectification module jointly corrects their image-plane locations and metric depths while estimating pointwise reliability. The rectified anchors are then selectively propagated before high-level multi-modal fusion, suppressing the influence of unreliable measurements. Experiments on ZJU-4DRadarCam and a newly collected 4D radar-camera-LiDAR dataset show that RbFT-Net consistently outperforms the evaluated independent radar-camera methods and remains competitive with plug-in pipelines using auxiliary monocular depth models. Cross-platform evaluation and component analyses further support the effectiveness of the proposed rectification and reliability-aware propagation strategy.
Wentao Zhao, Shouxuan Wu, Yongtao Cen et al.· 0 citations
Sensing ego-velocity estimation is fundamental to state estimation in visually degraded environments, where camera- and LiDAR-based pipelines can become unreliable. Millimetre-wave radar is well suited to these conditions because it provides direct Doppler velocity sensing and remains robust to poor illumination, textureless scenes, and airborne particulates. However, conventional radar ego-velocity pipelines typically apply constant false alarm rate (CFAR) thresholding to convert dense radar spectra into sparse point clouds, prematurely discarding sub-threshold returns that may still retain useful Doppler motion cues. We present Dense Soft Weighting, an analytic radar front-end that maps every range-Doppler cell to a continuous confidence metric rather than enforcing a binary detection threshold. Ego-velocity is then estimated using a deterministic robust weighted least-squares formulation, while the same weighted measurements provide a closed-form, measurement-derived velocity covariance for integration with a shared inertial back-end. The method requires no platform-specific training data or learning-based uncertainty model, supporting transfer across single-chip radar configurations. Across two public datasets and one self-collected dataset, Dense Soft Weighting reduces mean absolute pose error by 31-45% relative to the strongest CFAR point-cloud baseline under an identical inertial back-end, while running in real time on embedded hardware.
A. Babgei, Chenyu Zhao, M. Breza et al.· arXiv.org· 0 citations
Recent advances in 4D radar enable robust perception in adverse weather; however, the inherent sparsity, noise, and limited positional precision of radar point clouds pose significant challenges for registration-based odometry. In this letter, we propose RaDiVe, a 4D radar odometry framework designed to improve the accuracy and robustness of radar point-cloud registration. We introduce a distance-bounded Normal Distributions Transform (NDT), which improves optimization stability and computational efficiency by restricting the correspondence search to near-distance voxel pairs. To mitigate measurement ambiguity, we propose a velocity-discrepancy point uncertainty model that weights each input 4D radar point according to the discrepancy between its measured Doppler radial velocity and the radial velocity predicted from the estimated ego-velocity. Furthermore, we incorporate Signed Distance Function (SDF)-based surface point extraction via implicit neural mapping to construct a geometrically consistent and noise-filtered local submap. Evaluations across multiple public datasets demonstrate that RaDiVe outperforms existing 4D radar odometry baselines by 44.4% in translational Absolute Trajectory Error (ATE) and 21.3% in rotational ATE on average, while maintaining real-time performance. The source code will be made publicly available to the robotics community: https://github.com/to-be-open-sourced.
The proposed TD-PDA generalizes to unseen users with an ultralow inference latency, successfully reconstructing legible trajectories even in the presence of strong multipath interference and achieves stability comparable to a well-tuned classical PDA filter via a purely data-driven design.
Salah Abouzaid, Leander Nothelle, Nils Pohl· IEEE Transactions on Radar S...· 0 citations
The 3-D human pose estimation is a key task in the fields of computer vision and wireless sensing. Compared with optical sensors, through-wall (TW) radar can penetrate nonmetallic obstacles and capture reflected signals from targets, making it highly promising for applications in visually constrained environments. However, severe signal attenuation and low spatial resolution of TW radar make accurate human pose reconstruction still highly challenging. To this end, we propose an end-to-end multiperson 3-D pose estimation method based on multi-input multi-output (MIMO) TW radar and transformer (PERT). In PERT, we first extract fused multiscale features from horizontal and vertical radar heatmap sequences using a feature extractor. Subsequently, to enhance the focusing ability on effective regions in large-scale scenes and reduce computational overhead, we propose a signal-guided foreground selector (SFS) that leverages a signal-guided salience supervision mechanism to guide the selector in selecting tokens related to the targets. Finally, the spatiotemporal pose (STP) transformer module extracts fine-grained pose features from foreground tokens using an attention mechanism and predicts the positions of 3-D keypoints via a decoder. Experimental results demonstrate that PERT outperforms all baseline methods, achieving an average localization error of 6.20 cm in a large-scale $6\times 15$ m scene behind a 24-cm cement wall.
Suyun Sun, A. Kong, Jian Guo et al.· IEEE Transactions on Instrum...· 0 citations
Millimeter-wave (mmWave) radar enables privacy-preserving and illumination-robust human motion reconstruction, but training generalizable models typically requires costly paired radar-motion recordings. Simulation can scale such supervision, yet even physics-based simulators cannot fully reproduce real-world multipath, clutter, hardware-specific response statistics, or distance-dependent resolution degradation, leaving a sim-to-real gap. We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. To learn transferable signal and motion priors, we pretrain a multimodal radar encoder with a physics-informed domain-randomization curriculum designed to mitigate the sim-to-real gap by approximating real-world propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A dual-mode mapping module predicts either motion-code distributions for structurally constrained zero-shot reconstruction or continuous motion parameters for flexible adaptation from limited real data. We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that prevents any exact subject-environment-location-motion tuple from appearing in both the adaptation and test sets. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7-39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without fine-tuning.