Skip to content

Leveraging YOLOv8 and WiFi CSI for robust human activity recognition in indoor environments

Aug 2026 · Multimedia tools and applications · Vol 85 · 0 citations · 46 references

TL;DR

An innovative framework that integrates WiFi-based Channel State Information with the advanced object recognition power of YOLOv8 to enable robust, contactless activity classification and employs the deep learning capabilities of YOLOv8 for precise identification of diverse indoor actions.

View source

Similar papers

Review Open access Aug 2026

A Comprehensive Review of Wi-Fi-Based Indoor Human Activity Detection

This survey provides a systematic overview of cutting-edge research on Wi-Fi-enabled indoor human activity detection, classify mainstream technologies along three dimensions: signal processing pipelines, learning paradigms, and application granularity, and further dissect core challenges including environmental adaptability, data scarcity, and system scalability.

Zengqian Song, Jingming Li · 0 citations
Open access Aug 2026

Learning Shared Semantic Representations from WiFi CSI for Unified Multi-Task Human Activity Recognition

Human activity recognition (HAR) plays a critical role in intelligent wireless sensing and mobile edge computing. Compared with traditional vision-based and wearable-based approaches, WiFi channel state information (CSI) enables privacy-preserving and device-free activity perception. CSI encodes human-induced channel variations that serve as a natural basis for activity-oriented semantic analytics over wireless networks. However, existing WiFi CSI-based methods suffer from weak cross-scene generalization, high model complexity, and poor multi-task collaboration. To address these issues, this paper proposes UniSense-CSI, a unified multi-task framework that jointly learns dynamic gesture recognition, static posture classification, and fall detection. Specifically, the proposed framework converts CSI signals into pseudo-RGB images, extracts spatio-temporal features using a customized ConvNeXt backbone, and leverages an improved PerceiverIO module to compress high-dimensional features into a compact latent space. Based on the shared representation, task queries and adapters are introduced to enable parallel multi-task inference within a common architecture. Experiments on public datasets demonstrate that the proposed framework achieves accuracies of 99.64%, 99.66% and 96.09% on the three tasks, respectively, while maintaining low inference latency and favorable edge-deployment capability.

Jing-Lun Mao, Qi-Yue Ma, Fengxia Han · 0 citations
2026

MoCoNet: Motion-Aware Convolution for Wi-Fi-Based Multi-User Activity Recognition

Multi-user WiFi-based human activity recognition (HAR) with channel state information (CSI) is challenging because the received CSI contains overlapping motion-induced channel variations from multiple users, which complicates robust per-user activity inference. In this letter, we propose a motion-aware convolution framework that introduces signal-guided local aggregation for Multi-user WiFi CSI HAR. Compact motion cues are extracted from CSI phase and used to modulate convolution along temporal, subcarrier, and joint directions. This enables the model to emphasize coherent local CSI patterns within mixed multi-user observations, helping learn more discriminative and interference-aware CSI representations for multi-user HAR. Experiments on the WiMANS benchmark demonstrate that MoCoNet is an effective and complexity-balanced design, achieving 89.56% average accuracy under the standard environment-band evaluations, where it consistently outperforms representative baselines across the reported environment-band settings. In addition, under the 5 GHz leave-one-environment-out evaluation, MoCoNet achieves the highest average accuracy of 73.17% among the compared baselines.

Minh Tuan Pham, Phuoc Nguyen T. H. · 0 citations
Open access Aug 2026

A transformer-based framework for device-free human activity recognition using Wi-Fi CSI

Human activity recognition (HAR) plays a pivotal role in ambient assisted living, particularly for monitoring the elderly and patients with chronic conditions. However, traditional approaches relying on wearable sensors or video cameras face significant challenges regarding user compliance and privacy intrusion. To mitigate these issues, this paper proposes a device-free sensing (DFS) ( https://github.com/mestrelan/MDA-CSI ) framework utilizing Wi-Fi channel state information (CSI), named . We introduce a robust Transformer-based architecture designed to capture long-range temporal dependencies in wireless signals. was validated using a comprehensive dataset from 86 volunteers, ensuring high generalization capabilities across diverse human motion patterns.

Allan Costa Nascimento dos Santos, Pamella Soares, Iandra Galdino et al. · 0 citations
Conference 2026

LipHS : A Lightweight WiFi-enabled Human Sensing For Multi-Class Scenarios

WiFi-based human sensing technology utilizing Channel State Information (CSI) has garnered significant attention due to its reduced privacy concerns and the widespread availability of existing infrastructure, demonstrating broad development prospects in the field of intelligent computing. Deployment on edge devices represents its most prevalent application scenario. However, the high-complexity algorithms commonly employed to enhance sensing accuracy face substantial challenges when deployed on devices with limited computational resources. Furthermore, most existing studies conduct experiments only on datasets with a small number of categories. Although these approaches achieve high accuracy, they fail to meet practical sensing requirements. Consequently, developing high-accuracy, low-complexity, and practical WiFi-based human sensing systems remains considerably challenging. To construct an efficient and lightweight feature extraction network, we presents LipHS, a lightweight feature extraction framework capable of simultaneously capturing multi-level information from CSI signals. To further reduce the number of model parameters, we employ a channel pruning method based on Layer-Adaptive Magnitude-based Pruning (LAMP) scores. LipHS achieves model lightweighting while maintaining robust feature extraction capabilities. Experimental results demonstrate that the proposed LipHS method outperforms other baseline algorithms in sensing performance on complex multi-class gesture datasets.

ChunHao Xue · 0 citations
Preprint Aug 2026

WiFuse: An Attention Mechanism for Human Activity Recognition using Fused CSI Amplitude and Delay-Doppler Channel Features

Recently, Wi-Fi sensing has played a significant role in Human Activity Recognition (HAR), as it enables the detection of various activities using only Wi-Fi signals, ensuring privacy and remaining non-intrusive for the user. However, environmental characteristics such as reflective surfaces, hardware offsets, and other physical impairments affect recognition by the neural network, subsequently causing errors and significantly reducing model accuracy. To overcome this problem we present the WiFuse framework, a dual-stream Channel State Information (CSI) framework for human activity recognition (HAR) that pairs denoised time-domain amplitude variations with 2D-FFT-derived Delay-Doppler motion representations computed from the sanitized channel phase. The fused representation feeds a hybrid ResNet-Temporal Convolutional Network (TCN) neural architecture augmented with channel and spatio-temporal attention, where the ResNet extracts spatial-spectral features and the TCN models long-range temporal dependencies; a decoupled two-stage transfer learning strategy is employed to improve optimization stability and feature reuse. We conduct extensive experiments on two public datasets, including comparisons against state-of-the-art methods and alternative hybrid architectures, ablation studies, and cross-dataset and domain-adaptation evaluations. The proposed framework reaches an overall accuracy of up to 95.28% across the four environments of the XRF55 dataset and up to 98.20% on the multi-user Wi-MIR dataset. Overall, the results indicate that combining amplitude and Delay-Doppler representations within a dual-stream strategy, enhanced by transfer learning, improves recognition performance under conditions that typically degrade deep neural networks, such as class overlap, multipath propagation, noise, and interference.

Alison M. Fernandes, H. I. D. Monego, Bruno S. Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.