Experimental results confirm that IFMA-Stereo achieves state-of-the-art accuracy in remote sensing disparity estimation and effectively mitigates prediction errors caused by spatio-temporal heterogeneity, albeit at the cost of increased inference time compared to baseline methods.
Abstract
In optical remote sensing 3D reconstruction, high-resolution satellite stereo matching is a critical task, yet it is challenged by extreme imaging geometries, texture-less and repetitive patterns, occlusions, and scene variations caused by spatio-temporal heterogeneity. To address these issues, we propose IFMA-Stereo, an innovative binocular disparity estimation method that leverages a monocular depth foundation model. Our approach constructs a multi-scale spatial information pyramid to jointly integrate the foundation model with a disparity extraction network. At the feature level, an attention interaction mechanism captures multi-dimensional contextual dependencies and transforms general scene understanding priors into long-range associative features suitable for stereo cost volume construction. At the pixel level, a cyclic iterative refinement module embeds depth information from the foundation model throughout the iteration process and performs joint optimization, enhancing the model’s adaptability in geometrically complex regions. Experiments on the US3D and GaoFen-7 datasets demonstrate that IFMA-Stereo achieves superior performance in challenging areas (texture-less regions, disparity discontinuities, repetitive patterns) and effectively mitigates prediction errors caused by spatio-temporal heterogeneity, albeit at the cost of increased inference time compared to baseline methods. Quantitatively, the method achieves an end-point error (EPE) of 1.347 and a D1 error of 7.26% on the US3D dataset, and an EPE of 1.585 and a D1 error of 13.41% on the GaoFen-7 dataset. Notably, the method also yields precise predictions for unseen urban areas, indicating strong generalization. These results confirm that IFMA-Stereo achieves state-of-the-art accuracy in remote sensing disparity estimation.
The results highlight that careful parameterization — combining observation weighting, n-tuple point filtering, and per-satellite sensor refinement — is key to producing accurate, geometrically consistent large-scalemosaics from bi-satellite stereo imagery.
Michaël Erblang, Emelyne Saulnier, Guillaume Laurent et al.· The International Archives o...· 2 citations
Dense image matching is crucial in applications such as 3D reconstruction, autonomous driving, and remote sensing mapping; however, weak textures, occlusions, and large-disparity scenes remain challenging. To address these issues, this paper proposes a dense matching network based on a Transformer and multi-scale feature fusion, called Task-aware Multi-Scale Matching Network (TMSMNet). First, Swin Transformer is used to model global context in feature maps, enhancing the feature discriminability in weak texture regions. Then, a multi-scale cost volume is constructed, and adaptive fusion is achieved through deformable convolution to accommodate disparity variations of different ranges. Finally, an attention- guided iterative optimization module is introduced to improve the matching accuracy in occluded regions. Experimental results on the Scene Flow, KITTI-2015, and Middlebury datasets show that TMSMNet outperforms mainstream methods such as RAFT-Stereo on the D1-all metric of KITTI- 2015 and demonstrates good generalization and robustness. Ablation studies also confirm the effectiveness of each module. In summary, the method in this paper provides a feasible approach for dense matching. Future work will explore model lightweighting to support real-time applications and attempt to combine generative models to handle completely textureless regions, further enhancing its performance in complex scenes.
With the advance of deep neural networks, the quality of disparity maps obtained through stereo matching has steadily improved. However, existing stereo matching methods still struggle to preserve fine-grained geometric details, resulting in blurred edges and over-smoothed predictions in challenging regions. To address these limitations, we propose StereoDiffuer, an iterative diffusion-based stereo matching framework that explicitly models geometric details and progressively refines disparity estimates. The framework incorporates a Saliency Attention Perception (SAP) module to extract salient geometric cues, including object boundaries, thin structures, and sharp edges. Confidence-guided SAP features are combined with the initial disparity estimate to condition an iterative denoising diffusion process, which corrects residual disparity errors and restores geometric details suppressed during cost-volume regularization and upsampling. Experimental results on the Scene Flow and KITTI benchmarks demonstrate the effectiveness of the proposed framework and its competitive performance relative to the compared stereo matching methods.
Bohan Li· Signal processing. Image com...· 0 citations
LoRetta is proposed, a foundation model coupling matchability-aware affine localization with guided dense registration with LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels and establishes a unified evaluation protocol for sparse, semi-dense, and dense matchers.
Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produce reliable predictions in challenging regions, mainly due to the lack of strong geometric priors. In this paper, we propose $\textbf{GeoStereo}$, a unified stereo geometry estimation framework that leverages powerful diffusion priors to jointly predict disparity and surface normals. Specifically, GeoStereo couples a feed-forward stereo matching pipeline with a diffusion-based normal estimation branch. To enable effective interaction between the two tasks, we introduce a disparity to normal initialization strategy and construct a warp to left-view condition for the diffusion process. This coupled design allows the diffusion branch to provide strong structural priors that enhance disparity estimation in ill-posed regions, while the feed-forward branch offers reliable geometric guidance for accurate normal prediction. Extensive experiments show that GeoStereo performs reliably in challenging scenarios, including low-light environments, highly reflective surfaces, and transparent objects. Under zero-shot settings, it achieves Rank-1 disparity estimation on multiple benchmarks, including KITTI and NYUv2, and delivers the best normal estimation accuracy on many real indoor benchmarks, such as iBims-1 and ScanNet. Project page: https://qz-wei.github.io/GeoStereo.github.io/
Qizhe Wei, Xianda Guo, Shaocong Xu et al.· 0 citations
To address central pixel dependency and degraded matching reliability in low-texture or occluded regions, this paper proposes a robust binocular ranging framework integrating an improved Census-BT cost with a multi-resolution ROI pyramid strategy. The approach employs an annular Gaussian-weighted Census transform combined with weighted BT costs to alleviate the limitations of traditional Census transforms, significantly enhancing robustness against noise and complex environments. To optimize computational efficiency for real-time applications, an enhanced YOLO11 model is integrated to generate region of interest (ROI) layers for cross-scale disparity refinement, achieving a 2.8% improvement in precision compared to the baseline. This localization facilitates progressive disparity refinement across pyramid layers, enabling reliable depth estimation without the extensive supervised retraining characteristic of learning-based models. Evaluated on benchmarks, the algorithm demonstrates a favorable accuracy and efficiency trade-off compared to traditional non-learning baselines, achieving an 11.33% D1-all error rate on KITTI 2015 and a minimum mismatch rate of 9.52% on Middlebury 2021 under challenging conditions. System-level validation using a ZED stereo camera substantiates the framework’s practical efficacy. In real-world ranging experiments across diverse scenarios, the system maintains a maximum relative ranging error of 3.19% with practical measurement accuracy consistently exceeding 96.81%. With an execution time below 0.75 seconds, the proposed framework achieves an effective balance between accuracy and processing speed, making it suitable for robotic perception and industrial object measurement on resource constrained edge devices.
Unknown authors· Measurement science and tech...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.