Skip to content
Preprint

Self-Calibrating Dense Displacement Fields for Reliable Co-Registration of Large Optical Satellite Imagery

Aug 2026 · 0 citations · 41 references
Computer Science

Abstract

Co-registration underlies nearly every multi-temporal and multi-sensor use of optical satellite imagery, and operational products still carry documented offsets well above the fraction-of-a-pixel scale at which change detection, time series, and data fusion degrade. Real image pairs differ along several axes at once (sensor response, scene content, viewing geometry, resolution, mosaic seams), and the last of these is not a single global motion. Existing tools embed a motion model and constants tuned to their development data; a pair that fits is registered precisely, while one that does not either fails to match or returns a result wrong by tens of pixels with no failure reported. Learned matchers add a GPU requirement and carry no accuracy guarantee outside their training distribution. We present SCDF (self-calibrating displacement fields), a training-free, GPU-free estimator whose motion model is the dense per-pixel displacement field itself, so no scene motion falls outside the model. A single predict--measure--filter loop runs over a resolution pyramid: the accumulated field predicts where each patch of the moving image falls in the reference, RootSIFT matching and a correlation pass measure the displacement there to sub-pixel precision, and filters whose thresholds are all calibrated on the image pair itself decide what survives. One configuration, with no per-dataset tuning, processes full $8192^2$ scenes on a single CPU core. On 584 constructed-ground-truth pairs built from real Sentinel-2, Landsat-8/9, and NAIP imagery, against seven classical baselines and two zero-shot pretrained matchers, SCDF registers every pair with zero failures, reduces the best baseline's real-pair median end-point error from 6.83 to 4.17m, and cuts its 90th percentile from 17.8 to 7.77m.

View source

Similar papers

Preprint Aug 2026

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

LoRetta is proposed, a foundation model coupling matchability-aware affine localization with guided dense registration with LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels and establishes a unified evaluation protocol for sparse, semi-dense, and dense matchers.

Siwei Yu, Han Guo, Z. Shi et al. · 0 citations
Open access Sep 2026

A Geometry-Constrained Framework for Automatic Geometric Positioning Accuracy Assessment of Large-Scale Satellite Imagery

High-resolution optical satellite constellations continuously generate massive volumes of remote sensing imagery, making automatic, ground-control-point-free (GCP-free) geometric positioning accuracy assessment increasingly important for ensuring the quality of downstream applications. However, conventional GCP-free inspection methods based on local feature matching often exhibit limited robustness under large initial positioning errors, weak-texture regions, cloud contamination, and temporal appearance variations, resulting in poor generalization across large-scale production scenarios. To address these challenges, this paper proposes a geometry-constrained framework that integrates Rational Polynomial Coefficient (RPC) prior constraints, coarse-to-fine registration, adaptive match-density-based block selection, hierarchical geometric verification, and a geolocation residual confidence measure into a unified automatic quality inspection pipeline. The framework leverages LoFTR for dense feature matching, but its principal contribution lies in the system-level integration and operational design for large-scale industrial satellite image production. Extensive experiments on multi-satellite and multi-scene datasets from the Jilin-1 satellite series show that the proposed method achieves a median positioning error below 2 m, an Average Precision (AP) improvement of 0.53 over the baseline, and nearly perfect accuracy on the evaluated test set for confidence thresholds above 0.5. The framework has also been deployed in the operational production system of multiple commercial Jilin-1 missions for more than six months, demonstrating its effectiveness, robustness, scalability, and practical applicability for large-scale optical satellite imagery.

Jia-Ming Cui, Wei-Bin Wang, Li-Ming Fan et al. · 0 citations
Jul 2026

CorrelationFlow: A Training-Free Geometric Approach for LiDAR Scene Flow Estimation

LiDAR scene flow estimation has settled into a monoculture: nearly all recent methods share the same feed-forward architecture and the same family of self-supervised losses, inheriting each other's assumptions, and each other's blind spots. When those assumptions fail, as they do for sparse, distant, or fast-moving objects, every method built on them fails together, and adding parameters or simulated training data does not fix what the formulation itself gets wrong. This paper takes the opposite path. We present CorrelationFlow, a training-free geometric framework that reduces scene flow to two textbook operations: connected-component labeling and correlation maximization on bird's-eye-view occupancy images. Objects are isolated as spatio-temporal connected components, their motions recovered as correlation peaks, and the resulting velocities propagated to all member points. However, this dense correlation evaluates every candidate displacement of every cluster and requires a window of past sweeps; therefore, we develop a sparse counterpart that operates on a single sweep pair by matching lightweight occupancy descriptors at boundary key points. Because nothing is trained, nothing is inherited: on the multi-domain test set of the Argoverse 2 2026 Scene Flow Challenge, spanning five datasets with heterogeneous sensors and platforms, CorrelationFlow ranked second among unsupervised methods and degrades most gracefully at long range, where the shared assumptions of learned methods break down. Our results suggest that a substantial share of the scene flow problem is solvable by classical computer vision, and that progress may require questioning the formulation, not scaling it.

Minh-Quan Dao, Yancong Lin, J. S. B. Perez et al. · 0 citations
Preprint Sep 2026

Sub-Pixel Affine Registration of Space Debris Images via the Radon Point Spread Function

Inter-frame affine misalignment caused by platform jitter and attitude adjustments poses a fundamental challenge for multi-frame analysis of point targets in optical surveillance. Conventional registration methods rely on spatial intensity correlations or distinctive image features, both of which are largely absent in low-signal-to-noise-ratio point target imagery. We introduce the Radon Point Spread Function (RPSF) to characterize point targets in the Radon-transformed domain, and derive a closed-form framework that jointly estimates inter-frame translation and rotation from as few as four scalar RPSF samples per frame pair. The method requires no iterative optimization, feature extraction or interpolation, which is suitable for resource-constrained onboard processing. Simulation results confirm sub-pixel translation accuracy and a mean rotation error of 0.2556{\deg} at 1{\deg} Radon angular resolution. Validation on five real space debris datasets including both ground-based and in-orbit observations yields a mean calibration error below 0.5 pixels, substantially exceeding the precision required for reliable multi-frame processing.

Unknown authors · 0 citations
Open access Jul 2026

Global Block Adjustment for Mosaicked Stereoscopic Satellite Imagery

The results highlight that careful parameterization — combining observation weighting, n-tuple point filtering, and per-satellite sensor refinement — is key to producing accurate, geometrically consistent large-scalemosaics from bi-satellite stereo imagery.

Michaël Erblang, Emelyne Saulnier, Guillaume Laurent et al. · 2 citations
Review Open access Jul 2026

Bundle-Adjusted Initialization for Efficient Earth Observation Gaussian Splatting

Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.

Jiyong Kim, Shuang Song, Rongjun Qin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.