Skip to content

Cross-Modal Feature Adapter for Few-Shot Human Activity Recognition.

Jul 2026 · IEEE journal of biomedical and health informatics · Vol PP, pp. 1-14 · 0 citations
Medicine

TL;DR

A new cross-modal feature adapter is designed, which fine-tunes a pre-trained CLIP image encoder to effectively align image-sensor pairs and adaptively blend the old knowledge inherited from the original zero-shot CLIP with the new knowledge adapted from few-shot training samples, which makes training converges faster whilst forming a streamlined time series sensor encoder.

Abstract

Recent years have witnessed outstanding success of deep learning in sensor-based human activity recognition (HAR), spanning a wide range of real-world applications like healthcare management, fitness tracking, and fall detection. However, sensor data annotation scarcity still remains a main challenge unresolved, which requires human annotators to take a long-term and tedious observation to segment and timestamp sensor samples meticulously, hampering the wide use of deep learning models, especially in few-shot HAR scenario. To handle such issue, this paper introduces a cross-modal data augmentation, by exploiting activity label text as key words to search for activity-related images to construct an augmented dataset. On this basis, a new cross-modal feature adapter is designed, which fine-tunes a pre-trained CLIP image encoder to effectively align image-sensor pairs. Through a learnable residual ratio, it may adaptively blend the old knowledge inherited from the original zero-shot CLIP with the new knowledge adapted from few-shot training samples, which makes training converges faster whilst forming a streamlined time series sensor encoder. Extensive experiments and ablation studies are performed on three public HAR benchmarks. The experimental results demonstrate that the proposed method outperforms existing state-of-the-art HAR baselines under all few-shot scenarios. A practical on-device inference latency is provided.

View source

Similar papers

Jul 2026

Few-Shot Learning for Cross-Domain Human Activity Recognition Using Wearable Sensors.

A novel lightweight cross-domain few-shot sensor-based HAR network (CFSH-Net) is proposed for cross-domain activity recognition with limited labeled samples, which demonstrates strong cross-user generalization on PAMAP2 and USC-HAD, and stable cross-dataset transfer when trained on OPPORTUNITY and evaluated on four other datasets.

Hao Zheng, Hongji Xu, Fei Gao et al. · 0 citations
Jul 2026

A Meta-Learning Framework for Few-Shot Dynamic Hand Gesture Recognition with Soft Temporal-aware Contrastive Learning.

Few-shot learning improves data efficiency in surface electromyography (sEMG) gesture recognition, yet cross-subject and cross-posture generalization remains challenging due to nonstationary signals, electrode displacement, and time-varying within-gesture dynamics. We propose STC+Meta, a framework that couples temporally informed representation learning with efficient personalization. In pre-training, a soft temporal contrastive loss assigns adaptive weights by time lag to encourage temporal continuity while preserving short-lived discriminative transients, yielding coherent fused features from noninvasive sEMG and forearm accelerometer (ACC). In adaptation, implicit MAML provides data-efficient meta-updates that personalize the pretrained encoder to a new subject using only 10% of the subject-specific calibration data (∼28 s in our 7-class protocol), while matching the performance of subject-specific training that uses the full Session 1 data (∼19 min). On a self-collected multi-session dataset comprising eight upper-limb postures and seven gestures, the configuration that applies STC pre-training followed by multi-posture few-shot meta-adaptation achieves 97.63% accuracy when 10% of subject-specific data (∼28 sec) is used for calibration; in a 10-class task-transfer setting, it attains 95.75%. Relative to batch fine-tuning, STC+Meta shows faster early-stage convergence and lower inter-subject variability under limited calibration data. These results indicate that temporally aware contrastive pre-training combined with meta-learning enables calibration-efficient personalization for myoelectric gesture recognition under biased conditions.

Yu-Xiang Hu, Kunkun Zhao, Ziyu Cheng et al. · 0 citations
Jul 2026

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition

The proposed Joint Embedding Predictive Architecture framework designed to learn robust and generalizable representations from unlabeled datasets demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.

Mohd Halim Mohd Noor, AbdulRahman M. A. Baraka · 0 citations
Open access Jul 2026

Edgeefficient human activity recognition using a quantized patchbased transformer.

Human Activity Recognition (HAR) systems using deep learning have shown significant promise; however, deploying such models on edge computing systems remains challenging due to constraints in inference latency, memory footprints and computational capacity. This study proposes a lightweight, end-to-end patch-based Transformer architecture designed for efficient HAR for resource constrained edge environments, evaluated through an architectural ablation study and multiple quantization strategies using TensorFlow Lite. We conducted extensive experiments by varying patch lengths and Transformer encoder depths in order to evaluate the impact on classification accuracy, latency and model complexity. The results demonstrate that moderate temporal patching achieves an effective balance between temporal representation learning and computational efficiency. Experiments have been conducted on two widely open-source benchmark datasets, WISDM and PAMAP2, where raw sensor signals are pre-processed through segmentation, normalization, and overlapping windowing before being fed into a compact Transformer model with patch embedding and multi-head self-attention for feature extraction. The trained models are converted into FP32, FP16, INT8 dynamic range, and INT8 full formats to evaluate trade-offs between different metrics such as accuracy, model size, and inference latency. The baseline model achieves 96.77% accuracy on WISDM and 98.48% on PAMAP2, with consistently high macro F1-scores. Among all quantization variants, TFLite-INT8-Dynamic reduced model size by 95.55% and 89.80% for WISDM and PAMAP2, respectively and drastically improved inference latency by around 99% for both datasets with minimal accuracy degradation. These findings demonstrate that the proposed patch-based Transformer model for HAR has achieved an effective balance between recognition performance and computational efficiency, which indicates its strong deployment feasibility for resource constrained edge environments through TensorFlow Lite benchmarking and can provide a scalable solution for edge intelligence applications.

Aasif Rashid Khanday, Rajendra Kumar, Yonis Gulzar et al. · 0 citations
Open access 2026

A Semi-Supervised CPC-Transformer Approach for Human Activity Recognition Under Label Scarcity

A semi-supervised approach that integrates Contrastive Predictive Coding (CPC) with a hybrid BiGRU-Transformer architecture is introduced, thereby enabling comprehensive temporal modeling for human activity recognition in smart-home environments.

Ronak Fatahi, Fatemeh Sadat Lesani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.