Skip to content

CUTA-HAR: A Cross-User Temporal Attention Network for Wi-Fi CSI-Based Human Activity Recognition

Sep 2026 · IEEE Internet of Things Journal · Vol 13, pp. 40663-40674 · 0 citations · 55 references
Computer Science

Abstract

Wi-Fi-based human activity recognition (HAR) using channel state information (CSI) provides a nonintrusive and device-free sensing solution for smart cities, healthcare monitoring, and smart homes. However, recognition performance often degrades when models trained on limited users are applied to unseen users due to variations in body shape, posture, movement style, and surrounding conditions. To address this cross-user robustness challenge, this article proposes CUTA-HAR, a Cross-User Temporal Attention Network for Wi-Fi CSI-based HAR. CUTA-HAR combines multiuser supervised training with an attention-based bidirectional LSTM (BiLSTM) to capture informative temporal CSI patterns from multiple training users, without requiring data from the unseen test user during training. Experimental evaluations on a self-collected multiuser CSI dataset show that CUTA-HAR consistently outperforms representative sequence modeling baselines under a leave-one-user-out evaluation protocol, achieving average test accuracy improvements of 2.0%–6.7%. Action-level analysis further shows that structured activities can be recognized reliably, while complex activities such as fall and pickup remain challenging due to larger cross-user motion variations. These results indicate the effectiveness of attention-guided temporal modeling for improving cross-user robustness in Wi-Fi CSI-based HAR.

View source

Similar papers

Open access 2026

Low-Data Cross-Environment Transfer Learning for Wi-Fi CSI-Based Human Activity Recognition: An Inductive Bias Perspective

Wi-Fi channel state information (CSI)-based human activity recognition (HAR) has emerged as a promising device-free and privacy-preserving sensing approach. However, its practical deployment remains challenging because models trained in one environment often suffer substantial performance degradation when transferred to a different environment, particularly when only limited labeled target data are available. To address this issue, this study proposes a cross-environment transfer learning framework for CSI-based HAR that integrates CSI preprocessing, adaptive amplitude-phase fusion via TinyGate, an R(2+1)D backbone, and two temporal modeling strategies, namely Bidirectional Long Short-Term Memory (Bi-LSTM) and Transformer. The framework is evaluated on the MultiEnv dataset under both single-source and multi-source transfer settings using four target supervision ratios: 5%, 10%, 20%, and 40%. Additional validation is conducted using the Widar 3.0 dataset, and an additional Transformer configuration analysis is performed to examine the effect of depth, warm-up, and regularization. Experimental results show that Bi-LSTM generally achieves higher target accuracy, smaller source-target accuracy gaps, and more consistent adaptation behavior than the Transformer under the evaluated low-data transfer settings. In contrast, the Transformer requires more careful architectural and training design, including deeper architectures, stronger regularization, and larger target-label budgets, before its temporal modeling capacity can be translated into competitive transfer performance. These findings are consistent with the interpretation that the sequential inductive bias of Bi-LSTM is better aligned with the temporal continuity, structured noise, and environment-dependent variation of CSI signals. Overall, this research provides empirical evidence and practical guidance for designing Wi-Fi CSI-based HAR systems under limited-data cross-environment transfer scenarios.

Parma Hadi Rantelinggi, Mondher Bouazizi, Tomoaki Ohtsuki · 0 citations
2026

MoCoNet: Motion-Aware Convolution for Wi-Fi-Based Multi-User Activity Recognition

Multi-user WiFi-based human activity recognition (HAR) with channel state information (CSI) is challenging because the received CSI contains overlapping motion-induced channel variations from multiple users, which complicates robust per-user activity inference. In this letter, we propose a motion-aware convolution framework that introduces signal-guided local aggregation for Multi-user WiFi CSI HAR. Compact motion cues are extracted from CSI phase and used to modulate convolution along temporal, subcarrier, and joint directions. This enables the model to emphasize coherent local CSI patterns within mixed multi-user observations, helping learn more discriminative and interference-aware CSI representations for multi-user HAR. Experiments on the WiMANS benchmark demonstrate that MoCoNet is an effective and complexity-balanced design, achieving 89.56% average accuracy under the standard environment-band evaluations, where it consistently outperforms representative baselines across the reported environment-band settings. In addition, under the 5 GHz leave-one-environment-out evaluation, MoCoNet achieves the highest average accuracy of 73.17% among the compared baselines.

Minh Tuan Pham, Phuoc Nguyen T. H. · 0 citations
Jul 2026

Two-stream prototype network for Wi-Fi CSI-based cross-domain human activity recognition

Human activity recognition (HAR) using Wi-Fi channel state information (CSI) faces severe challenges in cross-domain generalization and data scarcity. Existing methods either rely on complex hardware deployment or suffer from insufficient spatiotemporal feature extraction, leading to poor performance under domain shifts. To address these issues, this paper proposes a lightweight two-stream feature extractor, Recurrent Convolutional Recursive Transformer—Multi-Scale Convolution Augmented Transformer (RCRT-MCAT), for few-shot cross-domain HAR. The model decouples CSI signals into a temporal stream and a channel stream to separately mine complementary spatiotemporal information. The MCAT branch employs multi-scale convolution and adaptive attention to capture fine-grained temporal patterns. The RCRT branch adopts recurrent convolution and recursive Transformer to efficiently model spatial dependencies across antennas and subcarriers. An evaluation framework is established on two public datasets, SignFi and Wiar, covering four experimental settings: in-domain recognition, cross-environment recognition, cross-environment cross-user recognition, and cross-dataset recognition. Experimental results demonstrate that, on the most challenging 76-way cross-environment gesture recognition task, the proposed model achieves an accuracy of 77.1% under the 1-shot setting, representing a 24.2 percentage point improvement over the FewSense model. When the number of samples is increased to 5-shot, the accuracy rises sharply to 91%, which is 28.2 percentage points higher than FewSense. In the cross-dataset recognition scenario, our model reaches a fine-tuned accuracy of 77.8% on user a2, 13.5% higher than FewSense. The average unfine-tuned accuracy across all users is 62.2%.

Xiaohong Huang, Yongzhi Xu, Kaiyue Zhang · 0 citations
Aug 2026

CGAC: A Convolutional Bidirectional GRU Network with Temporal Attention for WiFi CSI-Based Human Activity Recognition

WiFi channel state information (CSI) can characterize wireless-channel variations induced by human activities without directly capturing identifiable visual content, providing a contactless technical approach to indoor human activity recognition (HAR). To address the difficulty of a single convolutional or recurrent network in simultaneously modeling local fluctuations, long-range temporal dependencies, and key action segments, this paper proposes CGAC, a model that integrates convolutional bidirectional gated recurrent units with temporal attention. The model first uses one-dimensional convolution and max pooling to extract and compress local temporal CSI features, then employs a BiGRU to model bidirectional contextual dependencies, and finally applies single-vector temporal attention to adaptively weight key time steps. Multi-dataset evaluations are conducted on three public datasets: UT-HAR, NTU-Fi HAR, and NTU-Fi Human-ID. CGAC achieves an accuracy of 99.70% on UT-HAR and accuracies of 97.50% and 97.81% on NTU-Fi HAR and NTU-Fi Human-ID, respectively. The results show that CGAC delivers the best performance on UT-HAR and remains competitive across different acquisition tools and CSI classification tasks.

Lili Cai · 0 citations
Preprint Aug 2026

WiFuse: An Attention Mechanism for Human Activity Recognition using Fused CSI Amplitude and Delay-Doppler Channel Features

Recently, Wi-Fi sensing has played a significant role in Human Activity Recognition (HAR), as it enables the detection of various activities using only Wi-Fi signals, ensuring privacy and remaining non-intrusive for the user. However, environmental characteristics such as reflective surfaces, hardware offsets, and other physical impairments affect recognition by the neural network, subsequently causing errors and significantly reducing model accuracy. To overcome this problem we present the WiFuse framework, a dual-stream Channel State Information (CSI) framework for human activity recognition (HAR) that pairs denoised time-domain amplitude variations with 2D-FFT-derived Delay-Doppler motion representations computed from the sanitized channel phase. The fused representation feeds a hybrid ResNet-Temporal Convolutional Network (TCN) neural architecture augmented with channel and spatio-temporal attention, where the ResNet extracts spatial-spectral features and the TCN models long-range temporal dependencies; a decoupled two-stage transfer learning strategy is employed to improve optimization stability and feature reuse. We conduct extensive experiments on two public datasets, including comparisons against state-of-the-art methods and alternative hybrid architectures, ablation studies, and cross-dataset and domain-adaptation evaluations. The proposed framework reaches an overall accuracy of up to 95.28% across the four environments of the XRF55 dataset and up to 98.20% on the multi-user Wi-MIR dataset. Overall, the results indicate that combining amplitude and Delay-Doppler representations within a dual-stream strategy, enhanced by transfer learning, improves recognition performance under conditions that typically degrade deep neural networks, such as class overlap, multipath propagation, noise, and interference.

Alison M. Fernandes, H. I. D. Monego, Bruno S. Chang et al. · 0 citations
#explainable ai Preprint Aug 2026

XAI2CSI: Interpreting CSI with eXplainable AI for Human Activity Recognition

This paper introduces XAI2CSI, a framework that leverages eXplainable Artificial Intelligence (XAI) to analyze DL-based CSI sensing systems and employs SAGE, a model-agnostic explainability method, to quantify temporal, spectral, and spatial CSI contributions to HAR decisions under nominal and cross-context evaluations on IEEE 802.11ax data.

Idio Guarino, Alfredo Nascita, Domenico Ciuonzo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.