This work presents the first complete system for automated six degrees of freedom (6DOF) satellite pose estimation from spatially resolved, ground-based, adaptive optics (AO)-corrected imagery, addressing a key challenge in Space Domain Awareness (SDA). The approach mitigates the need for human labeling by directly regressing satellite orientation and position from blurry, noisy, and deeply shadowed imagery. A multi-stage deep neural network pipeline localizes the satellite, predicts pose, and optionally applies temporal filtering. Networks are trained exclusively on fully synthetic imagery generated from a CAD model, yet generalize effectively to real data, bridging the Sim2Real domain gap. On 137 real, human-labeled test images of Seasat, the model achieved a mean rotation error of 5° and a mean image-plane translation error of 21 cm. Slant range error was quantitatively evaluated on synthetic data due to unknown real-sensor parameters. Qualitative evaluation of additional real Seasat imagery rated 177 of 199 predicted poses as “ground truth equivalent” or “high-confidence match,” with zero catastrophic failures. The system was extended to seven degrees of freedom (7DOF) for satellites with articulating components and demonstrated on real Hubble Space Telescope (HST) imagery, achieving 5.5° rotation error, 51 cm image-plane translation error, and 8° symmetry-adjusted solar array error on a 249-frame pass with causal temporal filtering. Across 586 real test images from Seasat and HST (captured over multiple decades under diverse conditions) the system consistently performed well. Full 6DOF performance was quantified on a high-fidelity wave optics (HFWO) synthetic test set of Seasat, where the model achieved 8.4° mean rotation error, 34 cm image-plane translation error, and 1.4% line-of-sight range error at r0=6 cm and 1031 km range. In a limited 200-image benchmark, the model demonstrated 48% lower mean rotation error than a single human labeler while operating ∼800× faster. It required <40 h and a single A100 GPU to generate data and train. The approach was also demonstrated for ARGOS, a smaller satellite with highly symmetric geometry. An exploratory General Image-Quality Equation-based image quality metric (AO-IQ) was introduced as an empirical correlate for pose accuracy. General-purpose models like GPT-4o and Depth Anything V2 failed across most SDA tasks, but rapid gains in vision-language models warrant continued monitoring. These results establish a new operational baseline for practical, real-time satellite pose estimation from AO SDA imagery.
Accurate monocular pose estimation of noncooperative spacecraft is critical for autonomous proximity operations such as on-orbit servicing and active debris removal. A major practical challenge is the domain gap between synthetic training imagery and real Hardware-In-the-Loop (HIL) sensor data. We present a pose estimation pipeline built on a self-supervised DINOv3 Vision Transformer backbone with a deconvolutional heatmap head that localizes spacecraft keypoints. To bridge the synthetic-to-real gap without target-domain annotations, the pipeline combines multi-scale structural similarity (MS-SSIM) supervision, feature-level domain generalization, input-level style randomization, and iterative self-training with pseudo-labels from unlabeled HIL images. On the SPEED+ benchmark, the pipeline achieves state-of-the-art accuracy on both HIL domains using a single model with no adversarial training, no external augmentation networks, and no target-domain annotations.
Stefano Bergia, Irene Caracciolo, Fabrizio Stesina et al.· IEEE International Workshop...· 0 citations
High-resolution satellite imagery demands three-dimensional (3D) reconstruction methods that deliver both speed and geometric accuracy. Recent adaptations of 3D Gaussian splatting (3DGS) to satellite imagery demonstrate strong efficiency, but reconstruction quality often degrades under diverse illumination across multi-date, high-altitude acquisitions (with small intersection angles), limiting applicability to remote sensing and vision tasks. We present SatSplat, the first framework to adapt 2D Gaussian splatting (2DGS) to satellite photogrammetry, with online camera adjustment. We approximated satellite cameras with an affine model and learned a minimal delta parameterization for in-splat camera refinement from dense observations. The formulation was implemented with a 2DGS scene representation. To handle time-varying shadows and illumination changes, we integrated geometric shadow mapping and per-camera color correction during training. Across the evaluated DFC2019 and IARPA2016 benchmark sites, SatSplat achieved strong geometric accuracy while significantly outperforming prior 3DGS-based baselines. On our processed DFC2019 benchmark, SatSplat reduced mean absolute error by 11.93% and peak video memory by 31% relative to the previous state of the art. Our approach enabled large-scale digital surface modeling with practical computational efficiency. The project page is available at https://gdaosu.github.io/satsplat.
Shuang Song, Jiyong Kim, Rongjun Qin· Photogrammetric Engineering...· 1 citation
Abstract. Structure-from-Motion (SfM) pipelines rely heavily on the detection and matching of repeatable keypoints across images, yet the performance of modern learned feature extractors in challenging environments remains insufficiently understood. This paper evaluates classical and deep keypoint detectors for SfM reconstruction using winter Arctic UAV imagery, a domain characterized by low texture, repetitive patterns, and limited man-made structure. We compare three feature pipelines within a shared PyCOLMAP-based framework: SIFT with nearest-neighbor matching (SIFT+NN), SuperPoint, and DISK, along with a hybrid approach combining SuperPoint and DISK correspondences. Quantitative evaluation is conducted using standard SfM metrics, including number of observations, track length, observations per image, and reprojection error, complemented by qualitative analysis of keypoint distributions and reconstruction interpretability. Results show that SIFT+NN consistently achieves the most complete and stable reconstructions, producing the highest number of matched observations and lowest reprojection error across aggregate experiments. However, on more challenging subsets lacking clear structural features, learned methods demonstrate improved robustness, successfully reconstructing multiple views where SIFT fails. SuperPoint provides broader spatial coverage, while DISK produces denser clusters in high-confidence regions, highlighting complementary behaviors between learned approaches. Overall, the findings indicate that classical methods remain strong baselines for Arctic UAV photogrammetry under standard SfM pipelines, while learned detectors offer advantages in difficult conditions. The observed performance gap is attributed to domain mismatch and backend optimization for handcrafted features. These results suggest that domain-specific training and improved spatial feature distribution are promising directions for advancing learned keypoint methods in Arctic reconstruction tasks.
Nicholas Sansoterra, M. G. Lenzano, William J. Shuart et al.· The International Archives o...· 0 citations
Accurate satellite tracking requires up-to-date Two-Line Elements (TLEs), as outdated data can lead to significant positioning errors. While ground-based optical telescopes are highly accessible, generating TLEs from their data is complicated by the fundamental lack of direct range measurements. This paper presents an automated end-to-end pipeline designed to overcome this limitation by proposing a robust method to estimate the range from prior TLE. The pipeline can then generate the updated TLEs by calculating new satellite state vectors using the estimated range. The pipeline consists of star and satellite detection, astrometric calibration, and orbit determination. For star and satellite detection, the key component of the pipeline is a robust deep learning-based detection model. To achieve this, we benchmarked models such as Deformable DETR, RF-DETR, and YOLOv12 against traditional image processing methods with 3105 images in FITS (Flexible Image Transport System) format. RF-DETR yields the highest F1 score (0.93) and precision (0.97) at an 8-pixel threshold. The detected star coordinates resulting from using RF-DETR, the best-performing model, were fed into Astrometry.net for precise astrometric calibration to determine the satellite celestial coordinates. The range required for orbit determination was estimated by extracting the prior TLE and propagating it to the observation epoch via SGP4. The satellite state vectors were then calculated using TLE-constrained orbit determination approach using the estimated range, followed by an inverse SGP4 optimization to recover the mean orbital elements. The generated TLEs were validated against public TLEs from Space-Track.org. To evaluate this pipeline, updated TLEs were generated specifically for medium Earth orbit (MEO) and geostationary Earth orbit (GEO) targets. The results demonstrate that the proposed method yields high accuracy for GEO satellites, achieving a mean motion difference of 0.0030 rev/day and a 24-hour ground-track position error of 0.83 degrees. In comparison, MEO satellites achieve a mean motion difference of 0.0126 rev/day and an error of 5.27 degrees. These results suggest that the proposed pipeline provides a robust foundation for automated orbit determination with clear potential for further refinement.
Kamin Kanchanapradit, K. Noysena, R. Lipikorn· IEEE Access· 0 citations
Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.
Jiyong Kim, Shuang Song, Rongjun Qin· The International Archives o...· 0 citations
With the rapid growth of aerospace activities, space situational awareness (SSA) has become increasingly important for space security. Compared with conventional two-dimensional (2D) inverse synthetic aperture radar (ISAR) image sequences, three-dimensional (3D) representations provide richer structural information. In addition, novel view synthesis (NVS) supports continuous visual interpretation and helps compensate for observation gaps. However, the limited angular coverage of single-pass observations makes stable 3D reconstruction and high-quality NVS difficult without accurate geometric calibration. To address these challenges, Radar Neural Splatting (RNSplat) is proposed for 3D reconstruction and NVS from ISAR image sequences. Specifically, a Pose Head Adaptation via Reprojection (PHARE) module is introduced to refine viewpoint parameters under a cross-view reprojection consistency constraint, thereby improving the estimation of a 3D point map. Together with the associated geometric attributes, the estimated point map is then used to construct the 3D Gaussian splatting (3DGS) representation. To better reflect ISAR image formation characteristics, a Rendering Stabilization Unit (RSU) is further introduced, using a Gaussian-sinc kernel as a PSF-based scattering response approximation to improve cross-view synthesis quality. Experimental results demonstrate that the proposed framework improves NVS performance and enhances visible structural consistency. Ablation studies further show that PHARE improves geometric consistency and 3D point map quality, while RSU enhances the consistency and quality of synthesized views.
Huayong Tang, Guolin Ma, Dongcheng Li et al.· IEEE Transactions on Computa...· 0 citations
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.