Aug 2026· IEEE Transactions on Image Processing· Vol 35, pp. 8678-8690· 0 citations· 93 references
Computer ScienceMedicine
Abstract
Event-guided video super-resolution (VSR) leverages high-temporal-resolution event streams to address motion blur, rapid dynamics, and poor illumination that challenge frame-only VSR methods. However, most existing approaches emphasize reconstruction quality while overlooking real-time performance and computational efficiency, limiting their deployment in latency-sensitive scenarios. To overcome these issues, we present E2VSR, a lightweight and Efficient Event-guided VSR framework tailored for real-time applications. Operating under a causal setting with only current and past observations, E2VSR is designed for low-latency event-guided VSR. We propose an event-confidence adaptive propagation strategy comprising two key modules: the Event-induced Feature Modulation (EvFM) block for robust cross-modal event-frame integration, and the Event-Confidence Feature Fusion (EvCFF) block, which exploits events as motion cues for adaptive inter-frame aggregation. This design improves motion-aware temporal aggregation in challenging dynamic conditions, where event cues may provide complementary temporal information. Furthermore, an Implicit Event Reconstruction (IER) technique leverages event information during training to enrich feature representations without adding inference-time cost, enhancing spatial and temporal fidelity. Experimental results demonstrate that E2VSR achieves superior quantitative and qualitative performance while maintaining a low parameter count and computational cost.
This work proposes an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes and consistently outperforms existing state-of-the-art approaches.
Guixu Lin, Yuyang Yu, Xiang Ji et al.· 0 citations
Real-time 4K video super-resolution (VSR) requires the effective reuse of temporal information under strict latency constraints, typically relying on temporal alignment and real-time reconstruction. However, imperfect alignment introduces perturbations into the temporal recursion, which can accumulate over time and degrade reconstruction quality-an effect widely observed in recurrent VSR pipelines but rarely analyzed explicitly from a dynamical perspective under real-time constraints. In this work, we formulate real-time VSR inference as a recurrent dynamical process, embedding alignment within the state update. This formulation enables explicit analysis of how alignment-induced perturbations are introduced and propagated during inference. Motivated by this analysis, we propose StableFlow, a stability-guided real-time VSR framework. Specifically, StableFlow introduces an Alignment Gain-aware Alignment Module (AGAM) to enhance the utility of temporally aligned features, a State-aware Perturbation Control Filter (SPCF) to suppress unreliable propagated information, and a Temporal Propagation Control (TPC) loss to regulate long-term recurrent state evolution. Our method combines efficient temporal alignment with state-aware perturbation control and a temporal propagation control loss to stabilize long-term recurrent behavior. Experiments demonstrate that StableFlow achieves real-time 4K performance (over 40 FPS on an NVIDIA RTX 3090) for 4×upscaling from 960×540 inputs to 3840×2160 outputs, with only 321K parameters and 77.45G FLOPs, while maintaining competitive reconstruction quality. The source code will be made public after the peer review process.
Ke-Mi Chen, Xian-Bin Zhang, Ai-Ping Huang et al.· IEEE Transactions on Image P...· 0 citations
Event cameras offer microsecond-level temporal resolution and high dynamic range, but their spatial resolution remains much lower than that of modern RGB cameras. This paper studies high-resolution novel-view synthesis from low-resolution (i.e., low-spatial-resolution) event streams alone. Given multi-view low-resolution events of a static scene, without RGB images or any high-resolution signal, our goal is to reconstruct a 3D Gaussian radiance field that can be rendered beyond the native event-sensor resolution. To this end, we propose an event-only framework that integrates event super-resolution into event-driven 3D Gaussian optimization. The framework exploits two types of cues. Temporal cues convert the high temporal resolution of event streams into dense local multi-view constraints by constructing event observations between nearby viewpoints. Spatial cues provide target-resolution event priors by lifting low-resolution event increments with a 2D event super-resolution module. To make these priors compatible with the physical measurements, we apply pool correction so that each high-resolution prior reproduces the original low-resolution event increment after downsampling. The corrected high-resolution priors and the native low-resolution measurements are jointly used to optimize a shared 3D Gaussian radiance field, enforcing multi-view consistency during reconstruction. Experiments on synthetic multi-view scenes with paired low- and high-resolution event ground truth show that our method outperforms Pre-SR and Post-SR baselines in both quantitative metrics and visual quality, demonstrating the effectiveness of reconstructing high-resolution radiance fields from low-resolution events alone.
Event cameras are increasingly used for Multiple Object Tracking (MOT), but their asynchronous event output often requires specialized methods. Existing processing methods primarily follow two paradigms, pseudo-frames and event-by-event. The former is the prevailing approach since its data format aligns with images, making image-based techniques applicable. However, it suffers from tracking failures when trajectories overlap or are spatially close on pseudo-frames. Facing this challenge, we propose a multi-view pipeline, Multi-view Tracking (MvT), which preserves the 2D data format to leverage image-based techniques directly while introducing additional spatio-temporal views to resolve tracking ambiguities in a single view. MvT comprises a Multi-view Projection (MvP) module and a Multi-view Fusion (MvF) stage. MvP encodes events into three complementary spatio-temporal views while mitigating the pattern discretization. Within MvF, multi-view results are unified into a 3D coordinate system, and tracklets are associated through an optimization model subject to specific criteria combination. Evaluations on four datasets, including our self-collected Small Objects Dataset (SOD), show that MvT seamlessly integrates image-based methods and outperforms existing non-learning and learning trackers in generalized scenarios, and effectively resolves the single-view tracking ambiguities. Being training-free, MvT is applicable when ground-truth annotation is infeasible, thereby highlighting its practical, data-efficient potential. Code is available at https://github.com/zhazhabiu/MvTracking.
Muxi Zha, Banglei Guan, Minzu Liang et al.· IEEE Transactions on Image P...· 0 citations
ENCORE, an Event-Assisted Complementary Motion Refinement framework for learned video compression, employs Complementary Motion Representation to decompose aligned RGB-event features into common and modality-specific motion representations and identifies event-specific responses that are active and novel relative to RGB.
Shuhan Ye, Hong Yu, Chenqi Kong et al.· arXiv.org· 0 citations
Event cameras generate asynchronous, sparse data streams with microsecond temporal resolution, but in moderate-to-high motion scenes they can produce as many as hundreds of millions of events per second, creating significant bandwidth and storage challenges. Lossy compression is therefore essential for practical deployment, yet existing event stream distortion metrics fail to reliably predict compression-induced degradation at the task level, forcing codec optimization to rely on expensive task-specific evaluations. To address this gap, this paper introduces two fundamentally different event compression pipelines: i) an aggregation-based pipeline that converts the event stream into polarity-based histogram frames for compression with the conventional image codec JPEG 2000, and ii) a frame-free point cloud-based pipeline that codes events natively as 3D points using the octree-based codec G-PCC. Both pipelines are then assessed within a unified task-driven evaluation framework that relates event stream distortion to downstream application performance across four representative tasks: i) video reconstruction, ii) object detection, iii) optical flow estimation, and a delay-sensitive task iv) asynchronous feature tracking under a reference-relative protocol. Building on this framework, five classification-based distortion metrics are applied to event compression for the first time, to the best of the authors'knowledge, and benchmarked against existing event stream metrics. Experimental results demonstrate that the proposed metrics reliably predict compression-induced task degradation across different coding frameworks. This demonstrates that event stream distortion assessment can be an efficient alternative to repeated task-specific evaluation, providing direct guidance for the development and optimization of future event data coding solutions.
Zahra Rezaee, Catarina Brites, J. Ascenso· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.