Skip to content
#edge computing Open access

Adapting Pretrained Large Vision Models for Sensor-based Activity Recognition

Sep 2026 · Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies · Vol 10, pp. 1 - 31 · 0 citations · 131 references
Context-Aware Activity Recognition Systems

Abstract

Understanding and recognizing human activities from low-cost wearable sensors has attracted increasing attention in recent years. To achieve this goal, numerous learning models and augmentation approaches have been developed. While effective in certain scenarios, their performance is often limited due to the distribution shift between the training and testing data collected from different users or different body placements. Moreover, collecting large-scale and diverse sensor training data is costly and labor-intensive, which further exacerbates the data scarcity problem in HAR. To mitigate these challenges, in this work, we propose VisionHAR, a novel HAR framework that borrows knowledge from other data-rich modalities, i.e., low-cost and internet-scale image data. We first “draw” continuous sensor data on a figure to preserve both their temporal and periodic patterns. Then, we design a parameter-efficient transfer learning method to utilize the generalization capability of Large Vision Models (LVMs) pretrained on large-scale image data. To enable real-time activity recognition on edge devices, we further design a novel distillation approach to learn a highly effective and efficient student model, which achieves comparable performance with significantly fewer parameters. We evaluate our model in two main settings, i.e., cross-domain HAR and complex HAR (more than 18 activity categories). Experimental results show that VisionHAR outperforms the best existing HAR models by at least 8.22% in average accuracy and 11.48% in F1-score for cross-domain HAR with 1,460 times fewer parameters, and 7.62% in average accuracy and 7.86% in F1-score for complex HAR. Our findings suggest that the knowledge learned from vision data is generalizable to continuous sensor data, which provides a new potential solution for the data shortage issue in the activity recognition community. We release our code at https://github.com/saiketa/VisionHAR for future studies in the ubiquitous computing community.

Read PDF

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.