Skip to content

Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation

Sep 2026 · 1 citation · 104 references
Computer Science

TL;DR

This work presents an end-to-end pipeline that converts real-world driving footage from the Joint Attention in Autonomous Driving dataset into synthetic dynamic vision sensor (DVS) event streams using the v2e simulator, and trains a novel convolutional spiking neural network (Conv-SNN) with clip-consistent DVS augmentation to classify pedestrian crossing intent as binary: crossing or non-crossing.

Abstract

Anticipating whether a pedestrian will cross the road is safety-critical for autonomous vehicles, requiring real-time inference under challenging conditions including motion blur, high dynamic range, and class imbalance. Conventional frame-based deep networks process redundant RGB data at fixed frame rates, limiting their temporal resolution and energy efficiency. In this work we present an end-to-end pipeline that (i) converts real-world driving footage from the Joint Attention in Autonomous Driving (JAAD) dataset into synthetic dynamic vision sensor (DVS) event streams using the v2e simulator, (ii) augments training with the CARLA-simulated DVS sequences of the DVS-PedX dataset under both normal and adverse weather conditions, and (iii) trains a novel convolutional spiking neural network (Conv-SNN) with clip-consistent DVS augmentation to classify pedestrian crossing intent as binary: crossing or non-crossing. We detail all architectural decisions, the exact leaky-integrate-and-fire neuron dynamics with surrogate-gradient learning, the class-balanced loss formulation, JAAD oversampling at 6x, and a 70/15/15 stratified splitting protocol. The trained model achieves 95.83% accuracy and F1 = 0.9695 on the JAAD DVS test set, 97.79% accuracy and F1 = 0.9478 on normal CARLA DVS, and 94.78% accuracy and F1 = 0.8369 on adverse-weather CARLA DVS, all from a 1.07M-parameter architecture trained on CPU. Compared to prior frame-based approaches on JAAD, our method closes or surpasses the reported accuracy while operating natively on sparse temporal representations. We include a thorough analysis of the convergence behaviour across all 15 training epochs, domain transfer characteristics, and a quantitative comparison with representative related work.

View source

Similar papers

Open access Sep 2026

Real-Time Pedestrian Crossing Intent Prediction and Risk Assessment Framework Using Skeleton Graph Convolutional Networks

Pedestrian safety at urban intersections remains a major challenge in Intelligent Transportation Systems (ITSs). This study investigates whether crossing intention can be reliably inferred directly from temporal body-pose dynamics to drive real-time collision warnings on embedded edge platforms. Existing vision-based a...

Yi-Xuan Deng, C. Sub-r-pa, Rung-Ching Chen · 0 citations
#artificial intelligence Preprint Sep 2026

Unsupervised spiking feature learning for event-based pedestrian crossing detection: approaching supervised accuracy without labelled training data

The findings indicate that the accuracy cost of removing labels from feature learning is small on this benchmark, and that reported weaknesses of unsupervised spiking networks may be attributable to the readout protocol rather than to the learning rule.

Henok Teklu, Mustafa Sakhai, M. Mertik et al. · 0 citations
Open access Sep 2026

ADVANCED REAL-TIME TRAFFIC SIGN SEGMENTATION AND CLASSIFICATION USING HYBRID DEEP CONVOLUTIONAL ARCHITECTURES ON GTSDB AND BENCHMARK DATASETS

A rigorous comparative analysis of deep learning models, specifically U-Net, Mask R-CNN, and YOLOv8-Seg, for real-time traffic sign segmentation is presented, finding that YOLOv8-Seg achieves an optimal trade-off with a mean Average Precision.

Anaxon Muqimova · 0 citations
Aug 2026

Understanding and predicting rolling-gap pedestrian behavior under mixed traffic using pose-informed deep learning.

OBJECTIVES Pedestrian safety remains a critical concern in developing countries, particularly at unsignalized intersections characterized by non-lane-based, heterogeneous traffic. A prevalent and risky behavior in such contexts is rolling-gap crossing, where pedestrians opportunistically navigate through narrow moving-...

Kaliprasana Muduli, Indrajit Ghosh · 0 citations

A MICRO-HYDROPOWER PLANT

A rigorous comparative analysis of deep learning models, specifically U-Net, Mask R-CNN, and YOLOv8-Seg, for real-time traffic sign segmentation shows that YOLOv8-Seg achieves an optimal trade-off with a mean Average Precision and experimental results indicate that YOLOv8-Seg achieves an optimal trade-off.

Muqimova Anaxon · 0 citations
Open access Sep 2026

An Efficient Transformer Detector for Traffic Scenes via Lightweight Backbone and Multi-Scale Attention

Trans-DETR, an efficient end-to-end detector for traffic scenes based on a multi-scale attention mechanism and a lightweight backbone network, and a reparameterized detection head called RepHead, which deeply integrates cross-stage partial connections with reparameterization techniques, providing strong technical suppo...

Jin-Feng Zheng, Yuan Zhang, Yan-Bo Hui et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.