This survey presents a seven-category taxonomy of self-supervised learning methods covering: 1) input reconstruction or restoration, 2) context prediction, 3) contrastive learning, 4) feature clustering, 5) self-distillation-based feature reconstruction, 6) redundancy reduction, and 7) masked image modeling.
Abstract
The availability of large-scale labeled datasets has driven advances in AI-based computer vision, yet supervised learning remains costly and impractical in domains where annotation is scarce. Self-supervised learning (SSL) addresses this by harnessing unlabeled data to learn rich, transferable representations without explicit supervision. This survey presents a seven-category taxonomy of self-supervised learning methods covering: 1) input reconstruction or restoration, 2) context prediction, 3) contrastive learning, 4) feature clustering, 5) self-distillation-based feature reconstruction, 6) redundancy reduction, and 7) masked image modeling, with coverage extended to recent methods that include DINOv2, I-JEPA, SparK, data2vec 2.0, V-JEPA, DINOv3, V-JEPA 2, V-JEPA 2.1, C-JEPA, PhiNet v2. We situate this work within the existing survey landscape by explicitly comparing our contributions with prior SSL reviews. Beyond method descriptions, we provide: a chronological timeline of SSL evolution from 2008 to 2026; a cross-paradigm comparative analysis evaluating all seven families along collapse risk, scalability, computational cost, and downstream transferability; a dedicated comparative analysis of anti-collapse mechanisms; critical limitations and trade-off analyses per method family; and systematic benchmarking evidence on ImageNet-1K, PASCAL VOC, COCO, and five public medical imaging datasets. We also contribute a practical method selection decision matrix, extended challenge discussions, and actionable open problems for future research.
Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and visual structure. We propose a curriculum learning strategy for self-supervised Earth observation that ranks samples by geographic isolation, a label-free proxy derived entirely from geolocation metadata already present in geospatial datasets, requiring no image decoding, no model feedback, and no manual annotation. Unlike visual complexity proxies, it scales as O(D log D) with dataset size D and is well-defined for both contrastive and reconstructive objectives. We integrate the proposed measure into MoCoV2 and MAE pretraining and evaluate across three downstream tasks from CopernicusBench (BigEarthNet, DFC-2020, LCZ). Our curriculum reaches baseline final-epoch performance using as few as 20% of the training budget (MAE) and at most 40% (MoCo) of the training budget, and improves final downstream performance by up to +5 mAP on BigEarthNet, with gains of 1-5 points across benchmarks, matching visual-complexity curricula while reducing pre-computation cost by more than 140x (4 s vs. 568 s on SSL4EO). A CKA and effective-rank analysis further reveals that curriculum-trained encoders develop higher-dimensional, more uniformly utilized embedding spaces throughout training.
Daniele Rege Cambrin, Francesco Rossi, Mattia Varile· 0 citations
Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
M. L. Mekhalfi, M. M. Al Rahhal, Y. Bazi et al.· IEEE Geoscience and Remote S...· 0 citations
Data augmentation plays a central role in self-supervised learning, as the quality and diversity of augmented views strongly influence the learned representations. However, most existing self-supervised methods rely on fixed stochastic augmentation pipelines, while more adaptive alternatives often require expensive policy search, adversarial training, or additional optimization procedures. In this paper, we propose a lightweight learnable augmentation framework based on Extreme Learning Machines (ELM) for self-supervised visual representation learning. The proposed module predicts image-dependent transformation parameters and applies them through a differentiable augmentation operator, enabling joint optimization with the representation model while introducing minimal additional computational overhead. The framework is integrated into three representative self-supervised learning methods: SimCLR, BYOL, and SimSiam. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny ImageNet show that the proposed method consistently improves linear evaluation performance relative to reproduced baselines across most settings. In particular, the method yields notable gains on CIFAR datasets and remains effective on the more challenging Tiny ImageNet benchmark. A per-class difficulty analysis further shows that the proposed augmentation strategy substantially improves performance on hard classes, indicating stronger robustness to challenging categories while maintaining competitive overall performance. In general, the results demonstrate that lightweight learnable augmentation can effectively enhance self-supervised representation learning across different frameworks and datasets.
Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guidance.
Oliver Lemke, Alexander Liniger, Abel Gawel et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.