Aug 2026· IEEE Transactions on Image Processing· Vol PP, pp. 1-1· 0 citations
Medicine
TL;DR
A novel temporal iterative refinement (TIR) framework to obtain low-latency flow updates at high frequency with SNN-based feature encoders, and exhibits better cross-domain generalization, hinting toward the strong inductive biases of the network.
Abstract
Existing event-based optical flow approaches often build on frame-based counterparts, failing to deliver high-frequency flow estimation. Methods that specifically address this issue fail to achieve comparable performance or the desired computational efficiency. In this work, we introduce a novel temporal iterative refinement (TIR) framework to obtain low-latency flow updates at high frequency. The TIR module incorporates the previous flow estimate along with the updated feature maps to simultaneously update and refine the flow estimate at each time step, thereby predicting accurate nonlinear pixel trajectories. However, updating the feature space at high frequency with conventional CNNs may lead to the temporal aperture problem, as the small temporal receptive field may not be enough to capture the necessary spatial context. We introduce SNN-based feature encoders to efficiently address this problem. The temporal dynamics of the SNNs provide an increased temporal receptive field, while their deployment on neuromorphic hardware offers a promising path toward additional energy efficiency. The results obtained on the real-world MVSEC dataset show that our network achieves 17× and 33× lower computations than the state-of-the-art E-RAFT and TMA, respectively, while maintaining similar accuracy performance. Compared to other supervised learning-based approaches, our network exhibits better cross-domain generalization, hinting toward the strong inductive biases of the network. To demonstrate the remarkable potential of our approach, we also provide results in extremely challenging scenarios with highly nonlinear pixel trajectories from the MultiFlow dataset, which also features high-frequency ground truth. Our code will be available at https://github.com/AhmedHumais/STIRFlow.
A lightweight hybrid Spiking Neural Network and Convolutional Neural Network (SNN-CNN) computational imaging framework is proposed, providing a novel computational imaging approach and engineering paradigm for real-time quantitative optical measurement under extreme conditions.
Mai Zhang, Ze-Ren Gao, Yu Fu· International Conference on...· 0 citations
Motion deblurring is an essential capability for high-speed vision applications such as autonomous driving and unmanned aerial vehicles. Existing frame-based deblurring methods often struggle with rapid motion, whereas event cameras offer a promising alternative. However, irregular motion patterns and event noise remain challenging. To address these issues, we propose ALNet, an adaptive lateral-interaction spiking neural network for event-driven motion deblurring. ALNet adopts a dual-branch image–event architecture. In the event branch, we develop a Lateral Deformable Interaction Spiking Neural Network (LDSNN), which introduces lateral spike interactions and adapts the offset-and-modulation mechanisms of deformable convolution to the lateral interactions of spiking neurons. The design dynamically adjusts the interaction region, aiming to enhance motion boundary modeling and reduce sensitivity to event perturbations. For temporal feature extraction from event streams, we design an Image-Guided Spiking Transformer (IGST), which uses image-domain spatial context to guide the temporal processing of event spikes. For image–event feature fusion, we introduce a Cross-Modal Attention Fusion module (CMAF) to align and selectively fuse multi-modal features, thereby improving edge reconstruction. Experiments on the GoPro, REBlur, and Ev-REDS datasets show that ALNet contains only 6.08 M parameters while achieving competitive PSNR and SSIM performance, providing a favorable trade-off between restoration accuracy and model compactness.
Benefiting from high temporal resolution and dynamic range, event-based local feature methods have attracted increasing attention. However, event sparsity, noise, and limited texture still hinder robust local feature learning. Deploying such methods on resource-constrained platforms such as unmanned aerial vehicles also requires balancing accuracy and energy efficiency. To address these challenges, this paper proposes \textbf{E-S2Feat}, a spiking neural network framework for event-based local feature detection and description. The framework jointly optimizes local feature learning from the perspectives of feature representation and selection. First, a module-specific spiking activation mechanism preserves fine-grained structural cues and discriminative information under low-bit, energy-efficient inference, thereby improving overall representation fidelity. Furthermore, a semantic-guided feature modulation mechanism leverages semantic priors to refine keypoint response distributions and enhance local descriptor discriminability, thereby guiding the model to extract local features with greater geometric stability and stronger discriminative capability. Experiments on the ECD and EDS datasets show that the proposed method significantly outperforms baseline methods such as SuperEvent in pose estimation accuracy. It also achieves accuracy comparable to its artificial neural network counterpart while delivering an approximately 4.8-fold improvement in theoretical computational energy efficiency. Visual-inertial odometry experiments on the TUM-VIE dataset further verify the effectiveness and practical application potential of the proposed method in complete SLAM systems.
Yang Yi, Juntao Hua, Jinpu Zhang et al.· 0 citations
Autonomous systems require robust low-latency perception under rapidly changing scene dynamics and challenging illumination. In event cameras object detection commonly relies on recurrent architectures to accumulate sparse temporal information over time. This work investigates how temporal information can be encoded directly within the event representation. We propose a confidence-normalized continuous multi-timescale representation based on logarithmic B-spline temporal encoding together with a geometry-aware local confidence mechanism that exploits the spatial structure of event generation. Using a fixed feed-forward EventCenterNet detector, we show that the proposed representations consistently outperform the compact CSTR representation on PEDRo and Gen1 datasets. We further introduce a recursive exponential-polynomial approximation that enables efficient event-by-event updates while largely preserving detection performance. These results demonstrate that carefully designed event representations can capture a substantial portion of the temporal information learned through recurrent temporal modeling, providing a promising foundation for efficient feed-forward, event-driven, and future neuromorphic object detection.
Fredrik Lundell, Per-Erik Forssén, Mårten Wadenbäck et al.· 0 citations
Optical flow forms a fundamental information for various motion related vision problems: e.g., SLAM, visual odometry, and object motion estimation. Event cameras are ideal vision sensors for on-line, dynamic tasks that require optical flow estimation, as they have high temporal resolution, high dynamic range and low latency. However, efficiently decoding optical flow from events for high frequency operation while maintaining accuracy is still an open problem. Batch-based optical flow algorithms (CNN or contrast maximisation) accumulate event data over a short period of time and achieve state-of-the-art performance in terms of accuracy, but at the cost of algorithm latency and lower update rates (on par with traditional cameras). In contrast, event-by-event algorithms only compute flow vectors in small, local regions, achieving a lower latency, but losing accuracy when global information is ignored. In this paper, we introduce a spatio-temporal registration framework to increase accuracy of current state-of-the-art event-by-event flow estimation, while also introducing a twofold algorithm acceleration approach and a real-time implementation strategy to mitigate the impact of computation scaling with event rate. We evaluate our event by-event optical flow algorithm on MVSEC, achieving state-of-the-art results for event-by-event algorithms, and performance comparable to batch-based methods. Our method is also computationally efficient, enabling processing of the higher resolution DSEC dataset, and is the only event-by-event algorithm tested to run completely in real time. Furthermore, we demonstrate its effectiveness and efficiency through qualitative evaluations on the ECD and the high-resolution M3ED datasets. Finally, we introduce a moving object dataset, which is outside the autonomous driving domain, to evaluate the general applicability of the proposed optical flow algorithm. The code is available open-source [CODE AVAILABLE ON ACCEPTANCE].
Zhichao Li, Arren J. Glover, Lorenzo Natale et al.· IEEE Transactions on Pattern...· 0 citations
Multimodal temporal alignment is a critical task for applications such as audiovisual understanding, lip reading, and instruction following. However, in weakly supervised settings, challenges like asynchronous sampling, irregular event triggers, and coarse labels hinder precise cross-modal alignment. Frame-based methods rely on fixed time grids, leading to redundant computation in sparse-event scenarios and reducing event-level precision. Differentiable time-warping methods typically require high-resolution inputs, resulting in high computational costs and sensitivity to numerical parameters. To address these challenges, we propose SynNeura, an event-driven continuous-time liquid-spiking neural framework for fine-grained alignment under weak supervision. SynNeura models alignment as a continuous-time latent-state process, analytically propagating states between events and updating only when spikes occur. This results in computational complexity that scales with the number of events rather than the sequence length. SynNeura introduces a piecewise-analytic update using matrix exponentials and trace variables, along with a hierarchical contrastive alignment objective at spike, trajectory, and state levels, enhancing robustness and consistency. Experiments with the AVE, LRS2, and YouCook2 datasets show SynNeura consistently outperforms frame-based and continuous-time baselines in alignment accuracy, temporal consistency, and efficiency. SynNeura achieves 0.291 MAE and 0.713 CAS on AVE, 0.648 TC on LRS2, and 0.829/0.794 EP/ER on YouCook2, with an overall score of 0.738. These results show SynNeura is a scalable, interpretable, and efficient solution for event-driven multimodal temporal alignment.
Yu-Ping Zhang, Yan Liu· Neural Networks· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.