Skip to content

GeoFusion: Geometric Consensus Learning and Anchored Feature Fusion for Multi-View Depth Estimation

Sep 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 10641-10648 · 0 citations · 31 references

Abstract

Multi-view depth estimation is pivotal for 3D perception, particularly in collaborative scenarios requiring robust scene understanding. However, existing paradigms are often constrained by: (i) a reliance on photometric consistency for cost aggregation lacking explicit inter-view geometric verification, and (ii) symmetric fusion strategies where scale-ambiguous monocular priors can degrade the metrically grounded multi-view representations. To address these challenges, we present GeoFusion, a unified framework that enforces depth estimation performance through geometric consensus learning and anchored feature fusion. Specifically, we first introduce Geometric Consensus-Aware Cost Aggregation (GC-CA) to establish robust geometric consistency. By quantifying the deviation of per-view depth hypotheses from the global consensus, it explicitly identifies and suppresses inconsistent outliers. Furthermore, the Geometric-Anchored Feature Rectification Network (GAFR-Net) is proposed to implement directional feature fusion, where the multi-view cost volume serves as a geometric anchor to rectify monocular representations in metric space. This design effectively prevents scale-ambiguous monocular priors from corrupting the metric-accurate multi-view features. Extensive experiments demonstrate that GeoFusion achieves the best or tied-best performance on both the KITTI and 7-Scenes benchmarks, while exhibiting competitive zero-shot generalization on TUM-RGBD.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.