Skip to content
Open access

A Semi-Supervised CPC-Transformer Approach for Human Activity Recognition Under Label Scarcity

2026 · IEEE Access · Vol 14, pp. 121018-121042 · 0 citations · 43 references
Computer Science

TL;DR

A semi-supervised approach that integrates Contrastive Predictive Coding (CPC) with a hybrid BiGRU-Transformer architecture is introduced, thereby enabling comprehensive temporal modeling for human activity recognition in smart-home environments.

Abstract

Human Activity Recognition (HAR) in smart-home environments plays a vital role in applications such as ambient assisted living, elderly care, and health monitoring. Unlike vision-based HAR, smart-home systems rely on ambient sensors such as motion sensors, that capture sparse, asynchronous, and noisy data, making accurate recognition more challenging. However, the limited availability of labeled sequences in real-world homes poses a critical obstacle to traditional supervised learning methods. To address this limitation, we introduce a semi-supervised approach that integrates Contrastive Predictive Coding (CPC) with a hybrid BiGRU-Transformer architecture. CPC is utilized as a self-supervised pretraining stage to learn informative temporal representations from unlabeled sequences, which are then used by the downstream classifier. These representations are subsequently processed in parallel by Bi-GRU and Transformer components to model short-term and long-term temporal dependencies, respectively, thereby enabling comprehensive temporal modeling for human activity recognition. Experimental evaluations on two real-world environmental sensor datasets, CASAS Aruba and CASAS Milan, demonstrate that the proposed model outperforms several baseline architectures in semi-supervised settings, achieving improvements of 5.31 percentage points ( $\approx 6.4$ % relative) on Aruba and 9.37 percentage points ( $\approx 14.8$ % relative) on Milan.

Read PDF

Similar papers

Jul 2026

Few-Shot Learning for Cross-Domain Human Activity Recognition Using Wearable Sensors.

A novel lightweight cross-domain few-shot sensor-based HAR network (CFSH-Net) is proposed for cross-domain activity recognition with limited labeled samples, which demonstrates strong cross-user generalization on PAMAP2 and USC-HAD, and stable cross-dataset transfer when trained on OPPORTUNITY and evaluated on four other datasets.

Hao Zheng, Hongji Xu, Fei Gao et al. · 0 citations
Jul 2026

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition

The proposed Joint Embedding Predictive Architecture framework designed to learn robust and generalizable representations from unlabeled datasets demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.

Mohd Halim Mohd Noor, AbdulRahman M. A. Baraka · 0 citations
Open access Sep 2026

DMART-HAR: Dynamic Multimodal Transformer Learning for Cross-Domain Human Activity Recognition

Human activity recognition (HAR) in smart environments plays a critical role in applications such as healthcare monitoring, intelligent transportation systems, and ambient assisted living; however, existing approaches are limited by their inability to effectively handle heterogeneous multimodal sensor data, capture long-range temporal dependencies, and generalize across diverse real-world environments under domain shifts. In this work, we present DMART-HAR, a Dynamic Multimodal Activity Recognition Transformer framework that unifies structured multimodal representation learning, transformer-based temporal modeling, cross-modal interaction, and adversarial domain adaptation within a single architecture. Specifically, the developed method introduces a sensor tokenization mechanism to encode heterogeneous IoT data into a unified representation space, followed by a transformer encoder to capture global contextual dependencies, while a cross-modal attention module enables deep interaction among sensor modalities and an adversarial domain adaptation strategy enhances robustness to unseen environments. Extensive experiments on benchmark datasets, including CASAS, PAMAP2, and Opportunity, demonstrate that DMART-HAR consistently outperforms both conventional baselines and recent state-of-the-art methods, achieving accuracy/F1-scores of 94.3%/92.8%, 96.2%/94.7%, and 89.8%/88.1%, respectively, and consistently outperforms the strongest competing approaches under cross-domain evaluation settings. These findings demonstrate the effectiveness of modeling temporal dynamics, multimodal relationships, and domain invariance simultaneously, establishing DMART-HAR as a scalable and robust solution for real-world HAR applications.

Unknown authors · 0 citations
Open access Sep 2026

Self-Supervised IMU-Based Human Activity Recognition with Deep Spatio-Temporal Feature Extraction and Adaptive Feature Fusion

Self-Supervised Learning (SSL) has emerged as an effective paradigm for reducing the dependence of Human Activity Recognition (HAR) models on labeled data. To address the inadequate exploitation of IMU spatio-temporal correlations during pre-training and the limited generalization caused by simplistic fine-tuning strategies, a novel SSL framework for IMU-based HAR is proposed. The framework employs the Transformer and Depthwise Separable Convolution (DSC) to jointly capture global temporal dependencies and local spatial features, which are adaptively fused into discriminative spatio-temporal representations. These representations are subsequently enhanced through spatio-temporal feature extraction and multi-dimensional feature aggregation for downstream HAR. Furthermore, an IMU-based data acquisition platform was developed to construct the CQXY dataset. The proposed method was validated through comprehensive evaluations on four public datasets (UCI, Motion, HHAR, and Shoaib) and one self-collected dataset (CQXY). Experimental results show that, on the public datasets, the proposed method improves classification accuracy, F1-score, and Cohen’s kappa coefficient by an average of 13.11%, 14.24%, and 16.70%, respectively, compared with the baseline models. Similarly, on the self-collected dataset, the corresponding improvements reach 8.87%, 11.07%, and 10.81%. These results confirm the generalization of the proposed approach across datasets of different scales and domain.

Unknown authors · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.