Skip to content

MICA-Net: A Multimodal Cross-Attention Network for Human Action Recognition

Sep 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 61 references

TL;DR

A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.

Abstract

Automatic human action recognition (HAR) has become the most active research topic in recent years due to its broad applications, ranging from health monitoring and video surveillance to human-robot interaction. In this paper, we introduce a novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model. MICA-Net is composed of a lightweight 3D CNN, optimized for mobile devices, to extract visual features and a co-attention network that integrates a 2D CNN with a transformer to extract motion features. These extracted features are then continuously fused through dynamic gated Joint Cross-Attention Modules (JCAMs). These modules capture the intra- and inter- modal relationship while adaptively learning the contribution of each modality across different scenarios. We evaluated our recognition model on four publicly available multimodal datasets, MMAct, UESTC-MMEA-CL, UTD-MHAD, and MuWiGes. On MMAct, our model achieves an impressive F1-score of 89.24% with Cross-Subject and 96.57% with Cross-Session, outperforming current state-of-the-art methods. Similarly, on the UESTC-MMEA-CL, UTD-MHAD, and MuWiGes datasets, it achieves outstanding accuracy of 99.24%, 94.72%, and 98.98%, respectively. To demonstrate the practicality of the model, we design a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications. Real-time deployment indicates its potential for real implementation in the future. Our code is publicly available at https://anonymous.4open.science/r/MICA-Net-682E.

View source

Similar papers

Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning wit...

Şafak Kılıç · 0 citations
2026

Explainable 3D Convolutional Neural Networks Spatiotemporal Learning for Human Handshake Interaction Recognition

Human Activity Recognition (HAR) has gained significant attention in computer vision due to its wide range of applications in surveillance, social behaviour analysis, and human–computer interaction. Among various human-to-human interactions, handshake recognition is particularly important as it represents social intent...

S. Kumaravel, S. Veni · 0 citations
Conference Aug 2026

MSCALNet: a multiscale convolutional attention LSTM network for IMU-based human activity recognition

Wearable devices play an increasingly pivotal role in human activity recognition (HAR), particularly driven by the urgent demand in medical applications ranging from rehabilitation monitoring to fine-grained gait analysis. However, existing methods still struggle with insufficient exploration of cross-modal information...

Zi-Bo Wang, Runyang Lyu, Bin Xiao · 0 citations
Open access Aug 2026

Hand Gesture Recognition Based on Multi-Scale Attention Graph Convolutional Network

Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, a...

Xiaowei Han, Ting-Shan Yan, Yunjing Lu et al. · 0 citations
Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Open access Sep 2026

An open-set hybrid CNN-TCN network for human action recognition with unknown action rejection in collaborative robot workspaces

Introduction Open-set recognition capability is becoming necessary for collaborative robots because actual shop-floor dynamics do not always correspond to the fixed action classes used during training. Misclassification of unknown human motions may lead to inappropriate robot responses, whereas reliable rejection enabl...

Yasir Abdullah R, Ignisha Rajathi G, Barakkath Nisha U et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.