Skip to content

Similar papers

Open access Jul 2026

ERD: Extended RAW-Diffusion Framework for De-rendering sRGB Images

Abstract. Recovering RAW sensor measurements from sRGB images is a central problem in computational photography, as RAW data preserves the true scene radiance prior to the nonlinear transformations introduced by a camera’s Image Signal Processing (ISP) pipeline. However, ISPs vary across different camera brands and models, making inverse ISP reconstruction particularly challenging when the sensor characteristics are unknown. Existing approaches often rely on metadata, modifiable ISP, or camera-specific training, which limits their ability to generalize across unseen devices. In this paper, we investigate a diffusion-based inverse ISP framework designed for cross-sensor RAW reconstruction. Building upon the RAW-Diffusion model, we incorporate a ControlNet-guided architecture that provides structured conditioning to improve generalization without requiring metadata at inference time. Using the MIT-Adobe FiveK dataset, we evaluate seven camera models, which is sufficient to test cross-sensor robustness. Our results demonstrate that the proposed ControlNet-enhanced model enhances reconstruction accuracy on unseen sensors, outperforming the baseline RAW-Diffusion model on the Nikon dataset and achieving competitive performance on Canon, Leica, and Sony. These findings highlight the potential of guidance-based diffusion models for practical, camera-agnostic inverse ISP.

Jiaqi Shang, Yifan Qu, Jianbo Qi · 0 citations
Jul 2026

IDAM-RAW: a sensor-aware RAW imaging framework for low-light perception via illumination decoupling and adaptive feature modulation

Abstract. Low-light object detection remains fundamentally challenging due to the intrinsic misalignment between physical imaging characteristics and detection-oriented feature representations under extreme illumination degradation. This misalignment originates from the inconsistency between sensor-level signal formation and downstream representation learning, leading to unstable feature distributions and degraded detection performance. Existing approaches either rely on enhancement in the sRGB domain or directly learn from RAW data, yet both struggle to effectively bridge this gap. In this paper, we propose IDAM-RAW, a unified RAW-domain framework that bridges physical imaging processes and detection-oriented representation learning through task-driven end-to-end optimization. Specifically, DetISP maps RAW measurements to detection-friendly features without relying on fixed hardware ISP processing. The Residual Illumination Decoupling Module progressively reduces illumination-related variations in feature space and stabilizes optimization, whereas the Adaptive Feature Modulation Module suppresses interference propagation and enhances target-related responses across multiscale features. To support evaluation, we construct CR7-RAW, a real-world low-light bimodal dataset with spatially paired RAW and RGB observations, providing a new benchmark for RAW-based perception tasks. Extensive experiments on LOD, CR7-RAW, and BDD-Night demonstrate that IDAM-RAW achieves consistent performance gains across the evaluated datasets. Averaged over three independent random seeds, IDAM-RAW improves mAP@50 from 40.3 to 72.7 on LOD while maintaining consistent improvements on CR7-RAW and BDD-Night, supporting its effectiveness and cross-dataset robustness.

Min-Jie Dai, Xing-Yu Lai · 0 citations
Aug 2026

Low-light image enhancement technology based on vision transformer

This work validates the design effectiveness of decoupling global and local representations within a frozen backbone, and establishes a new baseline for parameter-efficient enhancement.

Yanpeng Cao, Yue Wang, Ming-Hui Liang et al. · 0 citations
Preprint Aug 2026

LiteKD-Net: Lightweight Knowledge-Distilled Network for Mobile Image Denoising

Mobile image denoising requires both good restoration quality and low computational cost. In addition, it's annoying to collect large-scale LQ-GT clean pairs. As a result, we propose LiteKD-Net, a lightweight knowledge-distilled network for mobile image denoising. First, a physics-guided noise simulation pipeline generates paired training data by adding pixel crosstalk compared with pipelines applied to cameras. Next, we adapt the Real-ESRGAN to identity-resolution denoising and construct a lightweight Student using Lite-RRDB blocks based on depthwise separable convolutions. Third, feature-level knowledge distillation is applied to transfer the Teacher's restoration capability to the Student without introducing additional inference cost. Experiments on real-world datasets show that our model reaches great reduction in runtime and increase in the inference rate with good restoration quality. Our model also reaches the best in all metrics compared with SwinIR. These results indicate that LiteKD-Net provides a great trade-off between restoration quality and computational efficiency.

Zhi-Yi Zhou · 0 citations
Open access Aug 2026

Binarized High-Efficiency RAW Video Restoration and Beyond.

This paper proposes BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation, and develops a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mitigate quantization errors.

Tianyu Zhu, Ying Fu, Hesong Li et al. · 0 citations
Open access Aug 2026

Deep Edge-Aware Post-Processing for JPEG Enhancement: CNN-Based Artifact Reduction and Image Quality Restoration

Joint photographic experts group (JPEG) is one of the most widely used image compression standards, but its lossy nature often introduces visible artifacts such as blocking, ringing, and blurring, particularly at lower quality factors. These degradations significantly reduce perceptual quality and affect downstream computer vision tasks. To address these limitations, in this study work a CNN-based edge-aware artifact reduction framework (CNN-AR) is proposed that integrates an enhanced deep super-resolution (EDSR) backbone with a holistically nested edge detection (HED) guided loss. This design enforces both pixel fidelity and edge consistency, enabling superior artifact suppression while preserving fine structural details. Extensive experiments conducted on benchmark datasets (LIVE1, Kodak, Set14, Classic5, and CLIC) across quality factors 10–40 demonstrate the effectiveness of the proposed approach. Compared to state-ofthe-art models including ARCNN, DnCNN, and DPW-SDNet, the proposed method consistently achieves higher perceptual scores. On average, CNN-AR improves PSNR by +0.38 dB, Structural Similarity Index (SSIM) by +0.012, MS-SSIM by +0.009, and PSNR-B by +0.41 dB across datasets, shows its ability to deliver both numerically superior and visually sharper reconstructions.

Nupur, Nishant Kumar, Sajal Suhane et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.