PortVE is a lightweight visual enhancement framework for adverse port weather conditions that combines degradation-aware Multi-Scale Pooling with Pool-Conv Downsampling and Frequency Directional Modulation and captures coarse-to-fine degradation cues, directional structures, and frequency-domain information while maintaining a compact computational profile.
Abstract
Vision-based perception systems are central to monitoring, inspection, and safety assurance in smart ports. However, adverse weather, particularly rain and fog, degrades image quality, suppresses structural details, and impairs downstream visual perception. Meanwhile, practical port vision systems are often deployed on resource-constrained edge platforms, imposing strict requirements on computational efficiency. To address these challenges, we propose PortVE, a lightweight visual enhancement framework for adverse port weather conditions. PortVE is built on an encoder-decoder architecture and combines degradation-aware Multi-Scale Pooling with Pool-Conv Downsampling and Frequency Directional Modulation. This design captures coarse-to-fine degradation cues, directional structures, and frequency-domain information while maintaining a compact computational profile. Experiments on public benchmarks and a self-collected port-scene synthetic weather dataset demonstrate that PortVE achieves strong restoration performance and efficient inference. It achieves 37.64 dB PSNR on SOTS outdoor and 34.21 dB PSNR on Test2800 with 6.17 M parameters and 42.83 GFLOPs. Downstream object detection experiments further demonstrate that PortVE improves detection robustness under adverse weather. The source code is publicly available at https://github.com/jssc-ZhaoBG/PortVE
Nighttime driving safety remains a critical challenge in modern transportation systems: insufficient ambient lighting significantly degrades visual perception quality, adversely affecting both human drivers and advanced driver-assistance systems (ADAS) and directly threatening road users’ safety. Traditional image enhancement methods often suffer from color distortion and visual artifacts, whereas existing deep learning approaches typically require paired training data and incur substantial computational overhead. To address these limitations, this paper presents BLEN (bio-inspired low-light enhancement network), a zero-reference deep learning framework that integrates biological vision principles with efficient convolutional architectures. Specifically, BLEN leverages Retinex theory for illumination–reflectance decomposition, is inspired by and functionally approximates lateral inhibition mechanisms for edge enhancement, and incorporates a Large-Kernel Convolution with Attention (LKCA) module that reduces the parameter count of the LKCA encoder block by 76% (0.56 M vs. 2.34 M for a standard 13 × 13 convolution) relative to standard large-kernel operations. Extensive experiments on the SICE and LOL benchmarks demonstrate that BLEN achieves state-of-the-art performance among real-time, edge-deployable zero-reference methods on the SICE benchmark, yielding a peak signal-to-noise ratio (PSNR) of 23.67 ± 0.14 dB and a structural similarity index measure (SSIM) of 0.891 ± 0.004 on SICE while maintaining 2.10 M parameters (2.1 MB in INT8, 8.4 MB in FP32). Furthermore, the proposed method enables real-time inference at 31 frames per second (FPS) on embedded platforms, including the HiSilicon SS928 and Jetson Nano, demonstrating that the proposed method is an efficient and effective front-end for camera-based ADAS perception on automotive-grade edge hardware.
GPE-YOLO is proposed, a robust detection framework built upon the YOLOv11 architecture that explicitly integrates multiscale edge priors to enhance feature resilience and validate the potential of GPE-YOLO for reliable deployment in real-world adverse weather scenarios.
Xiaojie Chen, Yi-Fei Zhou, Yi-Ming Zhou et al.· International Conference on...· 0 citations
This paper proposes BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation, and develops a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mitigate quantization errors.
Tianyu Zhu, Ying Fu, Hesong Li et al.· IEEE Transactions on Pattern...· 0 citations
The surging volume of high-resolution remote sensing (RS) images and the limited transmission capacity of the satellite-to-ground link impose a pressing challenge on image compression: how to maintain higher reconstruction fidelity at lower bit rates without compromising the reliability of downstream vision tasks (e.g., object detection). To address this challenge, we propose an end-to-end scalable remote sensing image compression (SRSIC) framework. Considering that downstream tasks prioritize semantic structure while visual interpretation requires textural details, we adopt a scalable framework to decouple these features. Specifically, the compressed bitstream is divided into a base layer and an enhancement layer. The base layer is dedicated to compact semantic features optimized for object detection via a feature transfer network, bypassing the need for complete decoding. The enhancement layer supplements residual details for high-fidelity image reconstruction. Furthermore, considering the complex scale variations characteristic of RS images, we design a multiscale asymmetric codec to extract multiscale features and employ an adaptive context entropy model to minimize redundancy. Experimental results on the DIOR dataset demonstrate that SRSIC achieves significant bitrate savings, while maintaining better image reconstruction quality and higher object detection accuracy.
R. Tang, Pei-Cheng Zhou, Jia Jia et al.· IEEE Journal of Selected Top...· 0 citations
Images captured under weak illumination often suffer from severe visual degradation, typically characterized by low brightness, reduced contrast, and obscured details. Consequently, effective enhancement of weakly illuminated images, such as those acquired in backlit or nighttime environments, is essential for reliable visual perception in applications including surveillance and autonomous navigation. Advanced software algorithms, such as multi‐scale pyramid fusion method, can achieve impressive enhancement performance, their high computational complexity and extensive nonlinear operations render them unsuitable for real‐time, low‐power embedded systems. To overcome these limitations, this paper presents a fully pipelined hardware implementation of a multi‐scale pyramid fusion enhancement algorithm for weakly illuminated images. The proposed architecture is implemented in Verilog and verified using the Vivado Simulator. By enabling high‐throughput, singleframe stream processing, the design effectively mitigates the latency and performance bottlenecks associated with softwarebased implementations, making it well suited for real‐time embedded vision applications.
Xiao-Xuan Wen, Han-Yang Ye, Xiao-Yu Ying et al.· SID Symposium Digest of Tech...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.