Skip to content

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition

Jul 2026 · arXiv.org · Vol abs/2607.16350 · 0 citations · 35 references
Computer Science Engineering

TL;DR

The proposed Joint Embedding Predictive Architecture framework designed to learn robust and generalizable representations from unlabeled datasets demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.

Abstract

Sensor-based human activity recognition (HAR) has achieved significant progressed in fully supervised learning settings. However, these supervised learning models rely on large amount of labeled data, which require labor-intensive collection and meticulous annotation. To address these challenges, this paper proposes a Joint Embedding Predictive Architecture framework tailored for sensor-based HAR, designed to learn robust and generalizable representations from unlabeled datasets. The proposed framework features an encoder designed to explicitly model both the fine-grained local temporal representations within individual window and the long-term temporal sequence of adjacent windows. Furthermore, we introduce an improved Variance-Invariance-Covariance Regularization (VICReg) objective function that incorporates computationally lightweight norm term to stabilize the JEPA pre-training phase. This term balances variance, invariance and covariance constraints to prevent representation collapse. The proposed HAR-JEPA framework is evaluated using two benchmark continuously performed activity datasets. The results show that high-quality representations are successfully learned by the proposed framework. Furthermore, the representations learned by HAR-JEPA demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.

View source

Similar papers

Open access Sep 2026

Self-Supervised IMU-Based Human Activity Recognition with Deep Spatio-Temporal Feature Extraction and Adaptive Feature Fusion

Self-Supervised Learning (SSL) has emerged as an effective paradigm for reducing the dependence of Human Activity Recognition (HAR) models on labeled data. To address the inadequate exploitation of IMU spatio-temporal correlations during pre-training and the limited generalization caused by simplistic fine-tuning strategies, a novel SSL framework for IMU-based HAR is proposed. The framework employs the Transformer and Depthwise Separable Convolution (DSC) to jointly capture global temporal dependencies and local spatial features, which are adaptively fused into discriminative spatio-temporal representations. These representations are subsequently enhanced through spatio-temporal feature extraction and multi-dimensional feature aggregation for downstream HAR. Furthermore, an IMU-based data acquisition platform was developed to construct the CQXY dataset. The proposed method was validated through comprehensive evaluations on four public datasets (UCI, Motion, HHAR, and Shoaib) and one self-collected dataset (CQXY). Experimental results show that, on the public datasets, the proposed method improves classification accuracy, F1-score, and Cohen’s kappa coefficient by an average of 13.11%, 14.24%, and 16.70%, respectively, compared with the baseline models. Similarly, on the self-collected dataset, the corresponding improvements reach 8.87%, 11.07%, and 10.81%. These results confirm the generalization of the proposed approach across datasets of different scales and domain.

Unknown authors · 0 citations
Open access 2026

A Semi-Supervised CPC-Transformer Approach for Human Activity Recognition Under Label Scarcity

A semi-supervised approach that integrates Contrastive Predictive Coding (CPC) with a hybrid BiGRU-Transformer architecture is introduced, thereby enabling comprehensive temporal modeling for human activity recognition in smart-home environments.

Ronak Fatahi, Fatemeh Sadat Lesani · 0 citations
Jul 2026

Few-Shot Learning for Cross-Domain Human Activity Recognition Using Wearable Sensors.

A novel lightweight cross-domain few-shot sensor-based HAR network (CFSH-Net) is proposed for cross-domain activity recognition with limited labeled samples, which demonstrates strong cross-user generalization on PAMAP2 and USC-HAD, and stable cross-dataset transfer when trained on OPPORTUNITY and evaluated on four other datasets.

Hao Zheng, Hongji Xu, Fei Gao et al. · 0 citations
Open access Jul 2026

Edgeefficient human activity recognition using a quantized patchbased transformer.

Human Activity Recognition (HAR) systems using deep learning have shown significant promise; however, deploying such models on edge computing systems remains challenging due to constraints in inference latency, memory footprints and computational capacity. This study proposes a lightweight, end-to-end patch-based Transformer architecture designed for efficient HAR for resource constrained edge environments, evaluated through an architectural ablation study and multiple quantization strategies using TensorFlow Lite. We conducted extensive experiments by varying patch lengths and Transformer encoder depths in order to evaluate the impact on classification accuracy, latency and model complexity. The results demonstrate that moderate temporal patching achieves an effective balance between temporal representation learning and computational efficiency. Experiments have been conducted on two widely open-source benchmark datasets, WISDM and PAMAP2, where raw sensor signals are pre-processed through segmentation, normalization, and overlapping windowing before being fed into a compact Transformer model with patch embedding and multi-head self-attention for feature extraction. The trained models are converted into FP32, FP16, INT8 dynamic range, and INT8 full formats to evaluate trade-offs between different metrics such as accuracy, model size, and inference latency. The baseline model achieves 96.77% accuracy on WISDM and 98.48% on PAMAP2, with consistently high macro F1-scores. Among all quantization variants, TFLite-INT8-Dynamic reduced model size by 95.55% and 89.80% for WISDM and PAMAP2, respectively and drastically improved inference latency by around 99% for both datasets with minimal accuracy degradation. These findings demonstrate that the proposed patch-based Transformer model for HAR has achieved an effective balance between recognition performance and computational efficiency, which indicates its strong deployment feasibility for resource constrained edge environments through TensorFlow Lite benchmarking and can provide a scalable solution for edge intelligence applications.

Aasif Rashid Khanday, Rajendra Kumar, Yonis Gulzar et al. · 0 citations
Sep 2026

Multimodal Action Recognition via Causality-Inspired Graph Representation Learning.

Multimodal human activity recognition (HAR) benefits from complementary skeleton, inertial, and visual observations. However, many learning-based models still treat relationships among modalities, joints, and sensor variables as symmetric associations. This limits their ability to represent asymmetric information flow and can weaken the preservation of modality-specific cues during feature fusion. We propose Causality-Inspired Structure Representation Learning (CSRL), a multimodal HAR framework that uses directional dependency modeling as a structural prior for representation learning. CSRL first estimates transfer-entropy-based graphs from temporal entities, including skeleton joints and IMU sensor variables. These graphs provide asymmetric priors that guide recognition-oriented graph learning in the representation space. CSRL further combines hybrid contrastive learning with an encoder-decoder architecture to learn modality-invariant, modality-specific, and structure-aware representations in a unified framework. This design encourages cross-modal alignment while retaining local motion cues that are important for fine-grained action discrimination. Experiments on five public HAR benchmarks, including UTD-MHAD, MMAct, CZU-MHAD, NTU RGB+D, and NTU RGB+D 120, show that CSRL consistently improves accuracy, F1 score, and recall over competitive supervised and contrastive baselines. These results support TE-guided directional structure modeling as a practical and interpretable prior for multimodal action recognition.

Yang Lan, Xin Long, Yao-Yuan Zeng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.