Skip to content
Preprint

Weather-Conditioned Depth Anything

Sep 2026 · 0 citations · 66 references
Computer Science

TL;DR

This work presents Weather-Conditioned Depth Anything (DA-W), a framework that explicitly disentangles style from content for weather-robust depth estimation, and introduces a Style Filter trained on a curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings.

Abstract

Monocular depth estimation foundation models, such as the Depth Anything series, have achieved remarkable performance across diverse domains. However, they still suffer from critical failures under adverse weather conditions, such as fog, rain, snow, or at night. To address this, we present Weather-Conditioned Depth Anything (DA-W), a framework that explicitly disentangles style from content for weather-robust depth estimation. Specifically, we introduce a Style Filter trained on a curated mix of real and synthetic degradation datasets to extract content-independent, degradation-aware weather embeddings. This style embedding is then injected into the Depth Anything backbone using a parameter-efficient, zero-initialized adapter. Such a lightweight modulation allows a single unified model to robustly adapt to diverse conditions, including fog, rain, snow, and low-light, while avoiding catastrophic forgetting of its core generalization abilities in normal conditions. We train the adapter using a pseudo-label distillation and alignment strategy. Our comprehensive experiments demonstrate that our proposed DA-W achieves state-of-the-art robust depth estimation, improving AbsRel by an average of 3.7% on our curated weather benchmarks, while matching or slightly outperforming performance on standard clean benchmarks. Our project page is available at https://zhaoming-tamu.github.io/WCDA/.

View source

Similar papers

Preprint Sep 2026

SFE-VGGT: Source-Free VGGT Distillation for Event-Based Monocular Depth Estimation

Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE...

Thai Duy Nguyen, Addison Lin Wang · 0 citations
Conference Sep 2026

Environment-aware dynamic prompting for self-supervised monocular depth estimation

Self-supervised monocular depth estimation (MDE) eliminates the reliance on expensive ground-truth depth annotations and has emerged as a powerful approach for a wide range of vision applications. However, current lightweight networks are hampered by two critical challenges: the limited representation capacity of stati...

Dong-Liang Wang, Ming Jin, Xiang-Qian Fang et al. · 0 citations
Conference Aug 2026

Monocular distance estimation: from geometric foundations and deep learning innovations to industrial deployment challenges

It is concluded that future industrial deployment on edge-computing platforms will rely on a synergy between lightweight network architectures and multi-sensor fusion and self-supervised frameworks.

Zi-Kang Fan, Zi-Hao Xiang, Jiang-Sheng Liu · 0 citations
#machine learning Preprint Sep 2026

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.

Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al. · 1 citation · ⚡1
#machine learning Preprint Sep 2026

The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators

Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in t...

Thomas Deixelberger, Markus Steinberger · 0 citations
Preprint Sep 2026

Real-World Perception for Autonomous Driving in Adverse Weather: Enhancing Standard Detectors via Foundation-Guided Auto-Annotation

Standard deployment-ready object detectors for autonomous vehicles degrade in adverse weather and lighting conditions without being trained on extensive domain-specific data. While large-scale vision foundation models offer robust zero-shot generalization, their high computational cost makes them impractical for real-t...

Sepideh Gohari, Goodarz Mehr, A. Eskandarian · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.