This work proposes GS-RealBlur, a data acquisition framework for real-world image deblurring, achieving both blur realism and acquisition flexibility, and introduces a Blur-aware Pose Refinement (BPR) module that refines the pose using appearance consistency and centroid alignment constraints.
Abstract
High-quality, large-scale paired data is essential for training learning-based image deblurring models. However, synthetic blurry images generally lack realism, while real-world captured images require complex and inflexible camera systems. In this work, we propose GS-RealBlur, a data acquisition framework for real-world image deblurring, achieving both blur realism and acquisition flexibility. Specifically, we use a handheld camera to capture blurry images, and deploy a gimbal to densely capture sharp images of the same scene. We reconstruct the 3D representation of sharp images and calibrate the camera pose of each blurry frame within this 3D. The image rendered from this 3D according to the pose serves as the sharp counterpart. To better align the rendered image with the blurry image, we introduce a Blur-aware Pose Refinement (BPR) module that refines the pose using appearance consistency and centroid alignment constraints. Leveraging GS-RealBlur, we construct a high-quality and diverse dataset. Extensive experiments demonstrate that a deblurring model trained on our dataset achieves superior generalization performance across various real-world deblurring benchmarks, consistently outperforming models trained on existing synthetic and real-world datasets. The code and dataset will be made publicly available.
An efficient generative framework designed to improve in-the-wild robustness under diverse real capture conditions and demonstrate strong perceptual quality, semantic fidelity, and temporal consistency on unseen videos, as well as improved robustness in downstream 3D reconstruction under severe motion blur.
Abstract. Recovering RAW sensor measurements from sRGB images is a central problem in computational photography, as RAW data preserves the true scene radiance prior to the nonlinear transformations introduced by a camera’s Image Signal Processing (ISP) pipeline. However, ISPs vary across different camera brands and models, making inverse ISP reconstruction particularly challenging when the sensor characteristics are unknown. Existing approaches often rely on metadata, modifiable ISP, or camera-specific training, which limits their ability to generalize across unseen devices. In this paper, we investigate a diffusion-based inverse ISP framework designed for cross-sensor RAW reconstruction. Building upon the RAW-Diffusion model, we incorporate a ControlNet-guided architecture that provides structured conditioning to improve generalization without requiring metadata at inference time. Using the MIT-Adobe FiveK dataset, we evaluate seven camera models, which is sufficient to test cross-sensor robustness. Our results demonstrate that the proposed ControlNet-enhanced model enhances reconstruction accuracy on unseen sensors, outperforming the baseline RAW-Diffusion model on the Nikon dataset and achieving competitive performance on Canon, Leica, and Sony. These findings highlight the potential of guidance-based diffusion models for practical, camera-agnostic inverse ISP.
Jiaqi Shang, Yifan Qu, Jianbo Qi· The International Archives o...· 0 citations
Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unposed images, they are not applicable when images are unordered or densely captured. In this work, we propose a method to train 3DGS from unposed images. Our method leverages a pre-trained 3D geometric foundation model as the neural scene representation. Since the accuracy of the predicted pointmaps does not suffice for accurate image registration and high-fidelity image rendering, we propose to mitigate the issue by initializing and fine-tuning the pre-trained model from a seed image. The images are then progressively registered and added to the training buffer, which is used to train the model further. We also propose to refine the camera poses and pointmaps by minimizing a point-to-camera ray consistency loss across multiple views. When evaluated on diverse challenging datasets, our method outperforms state-of-the-art pose-free NeRF/3DGS methods in terms of both camera pose
Yu Chen, Rolandos Alexandros Potamias, Evangelos Ververas et al.· Neural Information Processin...· 0 citations
Neural Radiance Fields (NeRF) achieves impressive novel view rendering performance by learning an implicit 3D representation from sparse view images. However, it is difficult to reconstruct a sharp NeRF from blurry input that often occurs in the wild. To solve this problem, we propose a novel Efficient Event-Enhanced NeRF (E3NeRF) framework, reconstructing a sharp NeRF by utilizing both blurry images and corresponding event streams. A blur rendering loss and an event rendering loss are introduced, which guide the NeRF training via modeling the physical image motion blur process and the event generation process, respectively. To improve the efficiency of the framework, we further leverage the latent spatial-temporal blur information in the event stream to evenly distribute training over temporal blur and focus training on spatial blur. Moreover, a camera pose estimation framework for real-world data is built with the guidance of the events, generalizing the method to more practical applications. Compared to previous image-based and event-based NeRF works, our framework makes more profound use of the internal relationship between events and images. Extensive experiments on both synthetic data and real-world data demonstrate that E3NeRF can effectively learn a sharp NeRF from blurry images, especially for high-speed non-uniform motion and low-light scenes.
Yunshan Qi, Jia Li, Yifan Zhao et al.· IEEE Transactions on Pattern...· 0 citations
Reconstructing editable Computer-Aided Design (CAD) models from images is essential for downstream modification, manufacturing, and design reuse. However, existing image-to-CAD methods are developed predominantly on synthetic renderings and face two coupled obstacles: a substantial appearance domain gap between synthetic and real images, and a previously overlooked parameter bias in widely used CAD data. We show that the local normalization adopted by DeepCAD concentrates several geometric parameters around a few discrete values while encoding substantial information in a single scale factor. Consequently, a model can achieve deceptively high parameter accuracy by exploiting these frequent values rather than inferring geometry from the input image. In this paper, we propose RealCAD, a unified framework that addresses these limitations at the representation, image, and feature levels. At the representation level, we redistribute scale information to the corresponding geometric parameters, producing less concentrated parameter distributions in a shared scale space. At the image level, geometry-constrained translation converts synthetic renderings toward the real-image domain while conditioning on object contours. At the feature level, a multi-positive contrastive objective aligns representations of the same CAD model across viewpoints and image domains, enabling CAD sequence prediction from each individual view. We further introduce OpenRealCAD, comprising four-view photographs of 392 3D-printed objects paired with ground-truth command sequences. Experiments show that the revised representation substantially reduces the accuracy attainable from parameter-frequency priors, making parameter accuracy a more reliable measure of image-conditioned geometric inference. RealCAD further improves real-domain command and parameter accuracy, while retaining competitive synthetic-domain performance.
Yi-He Sun, Zi-Yu Lu, Kaihua Tang et al.· 0 citations
This work presents FixAnything, a single model for fixing a wide range of rendering artifacts by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning.
Khiem Vuong, D. Ramanan, Srinivasa Narasimhan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.