Aug 2026· International Journal of Computer Vision· Vol 134· 0 citations· 91 references
TL;DR
This work proposes a novel multi-modal heat conduction operation framework for effective activity recognition, termed MMHCO-HAR, and conducts comprehensive experiments comparing the proposed method with baseline methods on the HARDVS 2.0 dataset and other public datasets.
The article compares the performance of traditional machine learning techniques with recent deep learning architectures such as CNNs, RNNs, TCNs, and Transformers, based on accuracy, computational cost, and suitability for real-world disorderly plotting.
Disha Deotale, Madhushi Verma, P. Suresh et al.· Discover Artificial Intellig...· 0 citations
Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic range, and low power consumption. While many studies and datasets in neuromorphic vision have focused on automotive and drone applications, human-centric daily-life scenarios remain largely underrepresented, despite their importance for developing and benchmarking event-based perception systems. Moreover, the few existing event-based human activity datasets are typically recorded with scripted human actions, limiting their ability to capture natural human behaviors. In this paper, we introduce EventKitchen, a large-scale stereo event camera benchmark dataset of human cooking activities in the kitchen. EventKitchen is egocentrically collected from 10 participants in 13 diverse kitchens, where the participants wear a helmet with multiple sensors and naturally perform cooking activities, without any scripted actions. EventKitchen comprises 5.5 hours of stereo event recordings with synchronized RGB, depth, and IMU data. We provide human annotations for 10,762 action segments and 13,482 bounding boxes. We train baseline models on EventKitchen to perform multiple event-based tasks, including action recognition, object detection, and stereo depth estimation. By capturing natural, real-world human activities, EventKitchen establishes a challenging benchmark for neuromorphic vision beyond autonomous driving.
Cheng Feng, Hesam Araghi, Liming Zheng et al.· 0 citations
Emotion analysis is a fundamental task in computer vision, but its practical deployment remains constrained by the privacy risks inherent to conventional RGB cameras. Bio-inspired event cameras present a promising hardware-level solution because they capture asynchronous brightness changes, thereby reducing exposure of facial identity details while leveraging high dynamic range for robust perception under challenging illumination conditions. Despite these advantages, existing event-based methods struggle in complex real-world settings due to limited dataset scales, simple acquisition conditions, and reliance on single-modality visual cues. To address these, we establish a challenging tri-modal benchmark with event, audio, and text modalities and propose the Information-Guided Gated Fusion (IGF) framework, which first pre-trains an event encoder on the FAU subset of Emo-DVS to capture fine-grained facial dynamics, then employs adaptive modality gating to suppress modality-specific noise, and finally leverages mutual information maximization to align robust cross-modal representations. To alleviate data scarcity, we introduce Emo-DVS, the first large-scale event-based emotion analysis dataset, which couples dynamic illumination with the Facial Action Unit (FAU) subset and emotion subset. Extensive experiments demonstrate that IGF achieves state-of-the-art performance.
Jia-Qi Chen, Qin-Fu Xu, Hao Zhuang et al.· 0 citations
Human Activity Recognition (HAR) is a fast-growing research area that focuses on identifying human actions using data collected from sensors and vision-based devices. It plays an important role in applications like health monitoring, smart homes, surveillance, sports analysis, and human-computer interaction. In recent years, several methods have been developed to improve the performance of HAR systems using machine learning, deep learning, and hybrid models. This paper presents a detailed review of different methods used in HAR. The study is divided into three main categories: vision-based methods, sensor-based methods, and hybrid approaches that combine both types. Each method is discussed with examples from recent research, along with their advantages and limitations. A comparison is also provided in the form of a table to highlight the performance and challenges of each approach. Although HAR systems have achieved good results in controlled environments, several challenges still remain. These include poor generalization to new users or unknown environments, difficulty in recognizing complex or overlapping activities, dependence on large datasets, and lack of real-time performance. This paper also discusses these research gaps based on recent findings. The future of HAR depends on building more accurate, reliable, and real-time systems that can adapt to different situations. The paper concludes by suggesting possible directions for future work, such as the development of lightweight models, use of standard datasets, better handling of real-time data, and making models more interpretable.
Human action recognition (HAR) plays a crucial role in safety monitoring, intelligent surveillance systems, and human-computer interaction applications. In this study, we evaluate and compare several deep learning architectures for HAR using the Weizmann dataset, using a YOLO-based preprocessing, including CNN, CNN with attention mechanism, MobileNetV2, and InceptionV3. The proposed YOLO-based preprocessing method was specifically designed to enhance feature extraction efficiency by isolating human subjects from background clutter, thereby reducing noise and improving spatial focus. Experimental results demonstrate that the YOLO-based CNN achieved state-of-the-art performance with an accuracy of 99.6%, significantly outperforming the CNN-Attention model (98.6%), MobileNetV2 (96.1%), and InceptionV3 (93.7%). These findings underscore the importance of robust preprocessing techniques and highlight the superiority of the proposed YOLO-based method in handling complex real-world scenarios.
M. Faris, Al Hakim, Regina Ayumi Ulayyaa et al.· 2026 7th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.