Event cameras can capture human motion with low latency and high dynamic range, providing an emerging sensing paradigm for action recognition. A large body of recent work converts event data into a sequence of dense frames and feeds them into a video classification model for prediction. Although the frame-based methods take advantage of pre-trained backbones from the image domain to achieve high accuracy, most existing works sacrifice the sparsity of events due to dense processing, thereby increasing model complexity and latency. To boost the efficiency of frame-based methods while fitting the sparse nature of events, we propose the Sparsity-Aware Vision Transformer Network (SVTNet) with two novel event-oriented designs. Instead of direct event-to-frame conversion, we introduce an event representation method named Progressive Cumulative Event Representation (PCER) that adaptively integrates spatiotemporal cues across multiple temporal slices to better preserve fine-grained information. To leverage the spatial sparsity of events to accelerate model inference, we present the Motion Prior-Guided Token Sparsification Module (MPTS) that prunes redundant visual tokens hierarchically based on the joint motion-semantic importance. Comprehensive experiments show that SVTNet achieves state-of-the-art accuracy on multiple benchmark datasets and maintains much lower model complexity than existing frame-based methods.
Bo-Cheng Xie, Jian Liu, You-Xuan Fang et al.· 2026 IEEE International Conf...· 0 citations
This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model's predictions.
Yuhan Liu, Linghui Fu, Zheng Yang et al.· Neural Information Processin...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.