Self-Supervised Vision Transformers for Spacecraft Pose Estimation Across Domain Shift
Abstract
Accurate monocular pose estimation of noncooperative spacecraft is critical for autonomous proximity operations such as on-orbit servicing and active debris removal. A major practical challenge is the domain gap between synthetic training imagery and real Hardware-In-the-Loop (HIL) sensor data. We present a pose estimation pipeline built on a self-supervised DINOv3 Vision Transformer backbone with a deconvolutional heatmap head that localizes spacecraft keypoints. To bridge the synthetic-to-real gap without target-domain annotations, the pipeline combines multi-scale structural similarity (MS-SSIM) supervision, feature-level domain generalization, input-level style randomization, and iterative self-training with pseudo-labels from unlabeled HIL images. On the SPEED+ benchmark, the pipeline achieves state-of-the-art accuracy on both HIL domains using a single model with no adversarial training, no external augmentation networks, and no target-domain annotations.