Skip to content
Open access

Lightweight mmWave radar human action recognition via knowledge distillation and physically enhanced representation

Jul 2026 · Measurement science and technology · Vol 37, pp. 336001 · 0 citations · 37 references
Physics

TL;DR

This paper introduces a physically enhanced input representation module that alleviates the representation limitations of sparse mmWave radar point clouds by explicitly incorporating inter-frame centroid differences and normalized timestamps as auxiliary motion and temporal cues and employs a high-accuracy dual-branch teacher model to guide a student model.

Abstract

Human activity recognition (HAR) based on millimeter-wave (mmWave) radar has attracted significant attention due to its advantages in non-contact sensing and privacy preservation. However, extracting robust fine-grained features from sparse and nonuniform mmWave point clouds remains challenging when relying solely on raw 3D coordinates, often leading to measurement uncertainty in complex dynamic scenarios. Furthermore, existing high-performance models incur high computational overhead, limiting their deployment on resource-constrained edge sensing devices. To address these challenges, this paper proposes a lightweight mmWave radar HAR framework based on knowledge distillation. First, we introduce a physically enhanced input representation module that alleviates the representation limitations of sparse mmWave radar point clouds by explicitly incorporating inter-frame centroid differences and normalized timestamps as auxiliary motion and temporal cues. Second, we employ a high-accuracy dual-branch teacher model to guide a student model that maintains macro-architectural consistency but utilizes simplified operators. A multi-granularity feature joint distillation strategy transfers representation capabilities to the student without incurring the teacher’s computational burden. Experimental results demonstrate that the student model achieves average recognition accuracies of 98.38% and 98.41% on the RadHAR and Pantomime datasets, respectively. Notably, the model requires only 0.27 M parameters and 0.029 GFLOPS.

Read PDF

Similar papers

Open access Sep 2026

Environmental Context-Aware Human Action Recognition from 4D Millimeter-Wave Radar Point Clouds

As an emerging sensing modality, 4D millimeter-wave radar provides a privacy-preserving and lighting-robust solution for indoor scene perception and human action recognition (HAR). However, existing radar-based HAR methods mainly focus on short-term motion dynamics while overlooking the influence of indoor scene context. Furthermore, reconstructing indoor layouts from sparse and irregular radar point clouds remains challenging. To address these issues, this paper proposes an environmental context-aware radar action recognition framework (ECRAR) that jointly models human motion and indoor scene context from 4D millimeter-wave radar point clouds. Specifically, a trajectory-to-floorplan inversion branch reconstructs indoor layouts from long-term trajectory point clouds by exploiting the relationship between human movement distributions and spatial accessibility. A motion feature extraction branch captures discriminative short-term motion representations through multi-scale spatial modeling, while an environmental context fusion module adaptively integrates scene and motion features using cross-attention. We also construct InActivity-Scene, a new dataset with fine-grained scene annotations and nine categories of daily and hazardous human actions collected in multiple indoor environments. Experimental results demonstrate that ECRAR consistently outperforms existing methods, achieving 92.2% recognition accuracy with significant improvements on environment-dependent actions, while ablation studies further verify the effectiveness of each proposed component.

Xiao-Zhi Li, Zhipeng Lou, Wei Liang et al. · 0 citations
Jul 2026

Wave2Body: Rethinking mmWave Human Pose Estimation as Radar-to-Body Token Translation

Millimeter-wave (mmWave) radar enables privacy-friendly human sensing, but its sparse point clouds are physical measurements of view-dependent electromagnetic reflections and only indirectly characterize body articulation. Recovering a complete 3D pose from such partial, geometry-dependent observations is therefore under-constrained. Existing methods directly regress joint coordinates from paired radar-pose data, relying on the same limited paired supervision to learn radar perception, human-body structure, and their alignment. This coupling can encourage dataset-specific shortcuts under ambiguous radar observations. We propose Wave2Body, a radar-to-body token translation framework that decouples these learning targets using a self-supervised mmWave tokenizer, a pretrained compositional body tokenizer that defines the output space, and a lightweight translator between them. Experiments on M4Human and mmBody show that Wave2Body achieves stronger cross-domain generalization than previous methods while incurring much lower computational costs for training and inference. All the code and experiment results are publicly available at https://github.com/Galaxywalk/Wave2Body.

Bo Liang, Chen Gong, Wei Gao et al. · 1 citation
Preprint Aug 2026

RaStream: Edge-Deployable Streaming Human Mesh Recovery from mmWave Radar

Millimeter-wave (mmWave) radar enables privacy-preserving human sensing for edge applications, but streaming SMPL-X recovery on edge devices requires accurate spatial evidence extraction and temporally stable predictions under lightweight causal inference. Sparse radar reflections make dense mesh recovery difficult, and heavy multi-scale spatial backbones can be costly for volumetric radar tensors while still diluting weak body evidence with background clutter. Frame-wise mesh estimates further exhibit jitter, while generic temporal models often mix slowly varying body morphology with fast pose and translation dynamics. We present RaStream, an edge-deployable radar-tensor streaming mesh recovery framework that combines a radar-aware spatial encoder with dual-state causal temporal refinement. The Radar-aware Spatial Structure (RaSS) encoder preserves 3D radar structure, localizes the subject, extracts body-centered evidence, and produces compact radar-aware tokens from short radar windows. The dual-state temporal module separates slow morphology state from fast motion state: it accumulates morphology evidence for shape and gender estimation through a token-conditioned update gate and tracks dynamic motion with a causal recurrent state. The resulting model keeps streaming memory fixed and avoids full-volume buffering. We formulate temporal sampling parameters $(T_w, T, s)$ that expose radar observation density, finite unroll horizon, warm-up/replay behavior, and output-rate tradeoffs, and evaluate reconstruction accuracy, temporal smoothness, and edge efficiency on M4Human. RaSS-Base reduces single-window MVE from 90.90 mm to 84.27 mm over RT-Mesh with fewer parameters, while RaStream further reduces MVE to 72.05 mm under the random-split protocol. Jetson Orin Nano profiling shows 26.93 ms FP32 latency for the Base configuration.

Jiazhen Dong, Lei Liu · 0 citations
2026

Through-Wall Multiperson 3-D Pose Estimation With MIMO Radar in Large-Scale Scenarios

The 3-D human pose estimation is a key task in the fields of computer vision and wireless sensing. Compared with optical sensors, through-wall (TW) radar can penetrate nonmetallic obstacles and capture reflected signals from targets, making it highly promising for applications in visually constrained environments. However, severe signal attenuation and low spatial resolution of TW radar make accurate human pose reconstruction still highly challenging. To this end, we propose an end-to-end multiperson 3-D pose estimation method based on multi-input multi-output (MIMO) TW radar and transformer (PERT). In PERT, we first extract fused multiscale features from horizontal and vertical radar heatmap sequences using a feature extractor. Subsequently, to enhance the focusing ability on effective regions in large-scale scenes and reduce computational overhead, we propose a signal-guided foreground selector (SFS) that leverages a signal-guided salience supervision mechanism to guide the selector in selecting tokens related to the targets. Finally, the spatiotemporal pose (STP) transformer module extracts fine-grained pose features from foreground tokens using an attention mechanism and predicts the positions of 3-D keypoints via a decoder. Experimental results demonstrate that PERT outperforms all baseline methods, achieving an average localization error of 6.20 cm in a large-scale $6\times 15$ m scene behind a 24-cm cement wall.

Suyun Sun, A. Kong, Jian Guo et al. · 0 citations
Open access 2026

Enhancing D-Band FMCW Radar Tracking for In-Air Writing Through Weakly Supervised Deep Association

The proposed TD-PDA generalizes to unseen users with an ultralow inference latency, successfully reconstructing legible trajectories even in the presence of strong multipath interference and achieves stability comparable to a well-tuned classical PDA filter via a purely data-driven design.

Salah Abouzaid, Leander Nothelle, Nils Pohl · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.