Self-Supervised IMU-Based Human Activity Recognition with Deep Spatio-Temporal Feature Extraction and Adaptive Feature Fusion
Abstract
Self-Supervised Learning (SSL) has emerged as an effective paradigm for reducing the dependence of Human Activity Recognition (HAR) models on labeled data. To address the inadequate exploitation of IMU spatio-temporal correlations during pre-training and the limited generalization caused by simplistic fine-tuning strategies, a novel SSL framework for IMU-based HAR is proposed. The framework employs the Transformer and Depthwise Separable Convolution (DSC) to jointly capture global temporal dependencies and local spatial features, which are adaptively fused into discriminative spatio-temporal representations. These representations are subsequently enhanced through spatio-temporal feature extraction and multi-dimensional feature aggregation for downstream HAR. Furthermore, an IMU-based data acquisition platform was developed to construct the CQXY dataset. The proposed method was validated through comprehensive evaluations on four public datasets (UCI, Motion, HHAR, and Shoaib) and one self-collected dataset (CQXY). Experimental results show that, on the public datasets, the proposed method improves classification accuracy, F1-score, and Cohen’s kappa coefficient by an average of 13.11%, 14.24%, and 16.70%, respectively, compared with the baseline models. Similarly, on the self-collected dataset, the corresponding improvements reach 8.87%, 11.07%, and 10.81%. These results confirm the generalization of the proposed approach across datasets of different scales and domain.