Generalized CSTR: Multi-Dimensional Representations for Event-Based Vision
Event-based cameras provide a powerful sensing modality for capturing dynamic scenes with high temporal resolution and low redundancy. However, leveraging modern deep learning architectures often benefits from transforming asynchronous event streams into suitable intermediate representations. The Compact Spatio-Temporal Representation (CSTR) addresses this by encoding events into a three-channel image-like format compatible with standard vision backbones, but its reliance on mean timestamps limits its ability to model long and complex temporal dynamics. In this study, we propose a Generalized Compact Spatio-Temporal Representation (gCSTR), a simple yet effective extension of the CSTR that enhances temporal expressiveness by constructing complementary representations in the spatio-temporal $xt$ and $yt$ projection planes. These representations preserve the fine-grained temporal structure while maintaining compatibility with conventional convolutional neural networks. To effectively combine multiple gCSTR representations, we introduce the Parallel Specialization Network (PaSNet), a multi-branch architecture. Each branch is trained on a distinct gCSTR view, allowing specialization to the structure of each representation. We evaluate gCSTR and PaSNet across a wide range of event-based object and action recognition benchmarks and demonstrate substantial improvements over the standard CSTR, particularly for long and complex action sequences. Our approach achieves state-of-the-art performance on several challenging benchmarks, while matching or closely approaching prior work on others, and provides new insights into when spatio-temporal projections are beneficial for event-based vision tasks.