The role of atmospheric scattering constraints in anchoring generative trajectories is analyzed to advocate physics-consistent, task-driven evaluation and develop unified hallucination benchmarks and efficient, physically constrained generative models for edge and safety-critical deployment.
Abstract
The optimization objective in image dehazing is fundamentally shifting from deterministic pixel mapping to high-dimensional probability distribution modeling. This survey organizes the field according to three optimization paradigms: physical prior-guided deterministic mapping, end-to-end feature reconstruction, and generative distribution alignment. Our synthesis yields three main findings. First, single-pass CNN, Transformer, and state-space models remain attractive for latency-sensitive applications, although pixel-wise objectives can suppress high-frequency and perceptually plausible details. Second, GAN- and diffusion-based methods often report improved perceptual or distributional quality when evaluated using LPIPS or FID, but iterative sampling and weak physical anchoring increase computational cost and the risk of structurally inconsistent details under dense haze. Third, heterogeneous datasets, image resolutions, evaluation protocols, and incomplete perceptual reporting do not currently support a controlled quantitative comparison of hallucination rates across architectures. We therefore analyze the role of atmospheric scattering constraints in anchoring generative trajectories and advocate physics-consistent, task-driven evaluation. Future work should develop unified hallucination benchmarks and efficient, physically constrained generative models for edge and safety-critical deployment.
GenRec is introduced, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow, and attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones.
Ata Çelen, Jaewoo Jung, Federico Tombari et al.· 0 citations
PixSDS is proposed, a lightweight VAE-consistent gradient repair method that decodes a latent SDS lookahead step and uses the decoded image as a clean direction for pixel-space optimization, reducing motion in VAE-inconsistent directions without retraining the diffusion model, changing the renderer, or replacing the SDS objective.
Fixed-view visual sensors require normal-image reconstruction that preserves structural detail while exposing a local deviation in a residual map. Standard denoising diffusion probabilistic models (DDPMs) use spatially uniform corruption. We study an image-derived, structure-aware noise schedule for reference-image reconstruction, where the input image is available and the same importance map can be fixed in the forward and reverse processes. A matched three-seed diagnostic further isolates the effect of the semantic-training/gradient-reverse map substitution and of a lower bound on local noise. Map consistency recovers much of the unconditional FID loss, but neither it nor the attenuation floor yields a universal advantage over DDPM. The six-category MVTec AD residual study is likewise category dependent: the gradient/edge special case improves selected texture-localization outcomes but degrades several object-level outcomes. We therefore present the method as a reconstruction diagnostic, not as a competitive industrial anomaly detector or a universally superior generator.
Xing-Yu Lu, Zeng-Shan Yao· Italian National Conference...· 0 citations
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web
Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al.· 0 citations
Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations. Recent approaches attempt to balance this tradeoff via posterior sampling or multi-stage generative pipelines, yet remain computationally expensive and architecturally complex. To overcome these limitations, we propose PCFlow (Perceptually Consistent Flow Matching), a unified framework that directly parameterizes a continuous transport from degraded observations to clean targets, jointly optimizing distortion and perceptual quality. While its latent consistency flow objective drives stable and efficient few-step inference, a Latent Consistency Perceptual Loss (LCPL) imposes semantic constraints directly on the guiding velocity field, steering the dynamics toward visually sharp data manifolds. Furthermore, recognizing the inherent conflict between structural and perceptual consistencies, we integrate a conflict-free gradient projection strategy to stabilize the multi-objective optimization landscape. Combined with lightweight, convolution-only backbone, PCFlow achieves competitive performance across diverse restoration tasks at a fraction of traditional computational costs.
Sangwoo Jo, Donggeun Ko, Jayeon Kang et al.· 0 citations
Monocular Depth Estimation (MDE) is one of the most rigorously studied problems in modern computer vision, yet it is fundamentally ill-posed. Recovering absolute three-dimensional geometry from a single two-dimensional projection is mathematically impossible without strong inductive priors. This paper presents a structurally organized review of MDE's evolution spanning from 2005 to 2026. We trace the trajectory from handcrafted Markov random fields to brute-force pixel-wise regression with convolutional neural networks (CNNs), and then to photometric self-supervision, which liberated the field from expensive LiDAR sensor suites. We further analyzed the vision transformers (ViTs) and generative diffusion priors have achieved unprecedented zero-shot metric generalization - with models such as Metric3D v2 attaining an Absolute Relative Error (AbsRel) as low as 0.039 on the KITTI benchmark. We also show that the standard self supervised baseline Monodepth2 degrades catastrophically to an AbsRel of 1.185 under nighttime conditions in NuScenes-Night dataset, which is a massive performance collapse from its clear-weather baseline. Physical-prior models such as PhysDepth recover this to 0.118 AbsRel by embedding Rayleigh scattering theory directly into the network. At the efficiency frontier, architectures such as LEDepth achieve 0.101 AbsRel at 5.7 ms inference time with only 3.1M parameters. We conclude that the future of depth estimation lies in embedding rigid physical and spatiotemporal priors into foundational latent spaces to ensure unyielding reliability in the physical world.
Hasan Mahmud Shanto, Mohammad Tofiqul Islam, Muhammad Ryan Hasan et al.· AIUB Journal of Science and...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.