Jul 2026· IEEE transactions on consumer electronics· Vol abs/2607.03013, pp. 1-1· 0 citations· 70 references
Computer Science
TL;DR
This work proposes MambaLIE, a Scene Light Intensity-Boosted Low-Light Image Enhancement method based on a State Space Model (SSM), which outperforms state-of-the-art CNN-based and Transformer-based LIE methods on four widely used synthetic benchmarks and five publicly available real-world benchmarks in terms of accuracy, speed, and model size.
Abstract
Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to sensor limitations and imaging pipelines, which degrades visual quality and affects downstream vision tasks. Existing methods based on Convolutional Neural Networks (CNNs) and Transformers have dominated current low-light image enhancement (LIE) due to their excellent ability to model hierarchical features. However, CNNs operate in local receptive fields that cannot model long-range dependencies, while Transformers overcome this problem but incur substantial computational costs. To address these challenges, we propose MambaLIE, a Scene Light Intensity-Boosted Low-Light Image Enhancement method based on a State Space Model (SSM). We first introduce scene light intensity to improve the structural distribution of illumination, which is then gated with the low-light input to guide enhancement. To better model the illumination while maintaining computational efficiency, we propose the Locally Enhanced State Space Model (LESSM) for efficient light enhancement. Our LESSM contains two branches: an SSM branch and a Local Enhanced branch, where the former is used to model the long-range dependencies with linear time complexity, while the latter is used to enhance local feature representations. Extensive experiments demonstrate that MambaLIE outperforms state-of-the-art CNN-based and Transformer-based LIE methods on four widely used synthetic benchmarks and five publicly available real-world benchmarks in terms of accuracy, speed, and model size, making it suitable for practical deployment on resource-constrained devices.
Results validate the effectiveness and robustness of the proposed illumination-aware modeling strategy for low-light image enhancement, IA2former, which effectively captures long-range dependencies, improves detail restoration, and preserves spatial structures under challenging illumination conditions.
Tianqi Jiang· Poster Volume 0007 The 2026...· 0 citations
Image enhancement is a widely researched area in the domain of computer vision, particularly image processing. Among the subdomains, low-light image enhancement (LLIE) receives considerable attention due to problems and challenges imposed by poor lighting conditions. As such, low-light images suffer from poor visibility, distorted colors, and loss of details, which limits their usability in many applications. The traditional methods have struggled to preserve such details and make the images susceptible to over-enhancement. Whereas, the learning-based techniques rely heavily on paired datasets for training. Therefore, we propose a Bidirectional Conv-GRU integrated GAN framework. Involving bidirectional Conv-GRU modules in our use-case enables the model to capture both local textures and long-range feature dependencies. Also, the use of unpaired datasets allows it to learn flexible and realistic mappings without the strict need for aligned image pairs. The results demonstrate the potency of our proposed work as compared to state-of-the-art methods.
Palak Deb Patra, Santosh Kumar Panda, Manoj Kumar Bishwal et al.· International Conference on...· 0 citations
Low-light image enhancement aims to improve visual visibility and perceptual quality under challenging illumination conditions. However, conventional convolutional neural networks (CNNs) are inherently limited in modeling long-range dependencies due to their restricted receptive fields, which often leads to insufficient global context modeling and suboptimal restoration results. To address this limitation, we propose MSHCDI-Net, a Multi-Scale Hybrid Cross-Domain Interaction Network that effectively integrates CNN and Transformer branches to jointly capture local texture details and global contextual relationships. Specifically, the proposed framework adopts a hierarchical encoder–decoder architecture to perform multi-scale feature extraction and progressive reconstruction. A cross-domain interaction mechanism is introduced to facilitate effective information exchange between convolutional and Transformer representations across multiple resolutions, enabling complementary modeling of fine-grained structures and long-range dependencies. Through adaptive feature fusion and multi-scale guidance, the network achieves improved structural consistency and detail restoration in low-light scenes. Extensive experiments on several public benchmarks demonstrate the effectiveness of the proposed method. MSHCDI-Net achieves 23.45 dB PSNR / 0.848 SSIM on LOL-v1, 23.74 dB / 0.910 SSIM on LOL-v2-synthetic, and 22.24 dB / 0.868 SSIM on LOL-v2-real, demonstrating competitive performance in both quantitative metrics and visual quality.
Bin Chen, Peitao Li, Chaobing Zheng et al.· PLoS ONE· 0 citations
This work validates the design effectiveness of decoupling global and local representations within a frozen backbone, and establishes a new baseline for parameter-efficient enhancement.
Yanpeng Cao, Yue Wang, Ming-Hui Liang et al.· Pattern Analysis and Applica...· 0 citations
This work proposes a model-driven deep neural network to effectively handle the joint degradation of low light and blur and designs an illumination enhancement module (IEM) and a reflectance refinement module (RRM) to improve brightness, restore fine details, and suppress noise.
Yao Xiao, You-Shen Xia, Zhen-Yu Lu et al.· IEEE Transactions on Neural...· 0 citations
Low-light image enhancement is crucial in situations where visible sensors might suffer from severe noise and information loss ( e.g., nighttime surveillance). Recent approaches investigate auxiliary modalities invariant to illumination to improve the performance, such as thermal infrared imaging. We propose a Multimodal Intrinsics-Guided Framework that integrates RGB and thermal data to reconstruct well-lit images. Our method utilizes a two-stage pipeline: first, we employ an intrinsic decomposition strategy to separate re-flectance and shading components through knowledge distillation, where a teacher network guides a student model in re-constructing consistent intrinsic components; then, a refine-ment stage restores fine structures and visual details. We train the proposed model on synthetic data from HDRT dataset and demonstrate strong generalization to real-world benchmarks such as LLVIP and V-TIEE, outperforming state-of-the-art methods in most evaluation metrics. Code is available at : https://github.com/simonemelc/TIRGlow
S. Melcarne, J. Dugelay· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.