Skip to content
Preprint

SSMB: Self-Supervised Local Feature Detection under Motion Blur

Aug 2026 · 0 citations · 54 references
Computer Science

TL;DR

SSMB is presented, a deblur-free, self-supervised keypoint detector for motion-blurred images that requires neither handcrafted detectors nor external pseudo-labels, and introduces the Local Discriminability Enhancement (LDE) module, which restores fine-grained local discriminability after global feature mixing.

Abstract

Keypoint detection under motion blur remains a significant challenge, as blur distorts local image structure and degrades the repeatability of feature localization. Existing approaches either rely on computationally expensive deblur-then-detect pipelines that may introduce restoration artifacts, or learn to regress the image positions of handcrafted keypoints extracted on sharp images, which reflects the assumptions of the handcrafted detector rather than what is truly repeatable under blur. We present SSMB, a deblur-free, self-supervised keypoint detector for motion-blurred images that requires neither handcrafted detectors nor external pseudo-labels. SSMB introduces the Local Discriminability Enhancement (LDE) module, which restores fine-grained local discriminability after global feature mixing. Training is performed in two stages. First, geometric pretraining on synthetic shapes bootstraps spatially discriminative keypoint detection without any external detector, just from the rendered geometry. Second, blur-aware training on real sharp-blur image pairs learns blur-invariant detection through a multi-component self-supervised objective that enforces cross-domain consistency, geometric alignment, and spatial coverage. Extensive evaluations on keypoint detection, image matching, relative pose estimation, and visual localization under motion blur demonstrate that SSMB establishes a new state-of-the-art among sparse keypoint detectors, consistently outperforming both supervised and self-supervised baselines across all tasks. Code, models, and datasets will be publicly available upon paper acceptance.

View source

Similar papers

Preprint Aug 2026

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

A generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries is proposed, and Sparse-Constraint Rectified Flow is introduced, a detector-oriented adaptation of Flow Matching for spatially sparse anomaly localization.

Jiangling Zhang, Shuxuan Gao, Zeyu Chen et al. · 0 citations
Open access Jul 2026

RA-LoFTR: Rotation-Equivariant Annular Convolution for Robust Detector-Free Matching

Image feature matching is a cornerstone of computer vision, yet robust correspondence estimation under rotation, low texture, and illumination variation remains challenging. Detector-free methods such as LoFTR reduce the dependence on repeatable keypoints, but their standard convolutional backbones are still sensitive to orientation changes and weak local structures. To address these limitations, we propose RA-LoFTR, which enhances LoFTR with Rotational Coordinate Convolution (RCC) and Adaptive Fusion Weighting (AFW). RCC partitions feature neighborhoods into concentric annular regions and aligns dominant orientations through channel-wise cyclic shifts, transforming rotation handling from discrete angle classification into channel-phase alignment. AFW dynamically fuses RCC-derived local structural cues with Transformer-based global positional information, and MAGSAC++ is further used as a geometric verification step for outlier rejection. On MegaDepth, RA-LoFTR achieves pose-estimation AUCs of 37.4, 53.6, and 66.0 at 5°, 10°, and 20°, improving over LoFTR by 1.1, 1.7, and 1.3 points, respectively. On HPatches, it improves homography AUCs over LoFTR by 2.2, 1.5, and 1.9 points at 3, 5, and 10 pixels, respectively. Robustness experiments under low-light, motion blur, Gaussian noise, and low-texture conditions further validate the proposed annular convolution design.

Xiangjin Zeng, Lihang Chen, Fan Fu · 0 citations
Conference Sep 2026

Understanding Domain-Shift Immunity in Deep Deformable Registration

Deep learning has achieved remarkable success in deformable image registration, yet the visual information that drives deformation estimation remains poorly understood. Rather than pursuing incremental performance improvements, this work investigates the fundamental source of robustness in deep registration models. Using diverse, domain-agnostic synthetic datasets, we decouple deformation learning from application-specific appearance and show that domain-shift immunity is an inherent, largely architecture-agnostic property of deep de-formable registration when trained with a robust pipeline. To identify the mechanism underlying this immunity, we compare models trained directly on raw image intensities with models operating exclusively on local feature representations extracted by a fixed, pre-defined feature extractor. The comparable performance of these models provides strong empirical evidence that deformation estimation is governed primarily by local structural features, rather than global, domain-specific appearance cues. These findings offer a principled explanation for the cross-domain generalizability of deep registration networks and point toward feature-centric designs for domain-independent registration.

Mingzhen Shao, Sarang C. Joshi · 0 citations
Jul 2026

GlobalForge: Towards Robust AI-Generated Image Detection

The proposed GlobalForge improves average BAcc on 8 in-the-wild benchmark groups by $\mathbf{5.89\%}$ over the previous state-of-the-art, and is clearly ahead of representative baselines on RealDeg-Bench under both single and compound degradations.

Manni Cui, Ruiqi Liu, Dianyuan Zou et al. · 0 citations
Conference Jul 2026

HASO-DETR: hybrid attention small object detection based on RT-DETR

The proposed framework features a redesigned cross-scale feature fusion module, CCFM-S2, which utilizes the SPD-Conv operator for information preserving downsampling and explicitly integrates high-resolution shallow features (S2 layer), thereby infusing indispensable spatial details into the feature hierarchy for small targets.

Yi-Fei Zhou, Xiaojie Chen, Yi-Ming Zhou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.