Skip to content

UvA-DARE (Digital Academic Repository) Elastic ViTs from Pretrained Models without Retraining

· 0 citations · 79 references

TL;DR

SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retraining or labels.

View source

Similar papers

Preprint Aug 2026

LeVJEPA: Efficient&Scalable Video Pretraining without the Heuristics

Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coupling an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor, or circumvent it by reconstructing masked content in pixel space. We introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective, which dispenses with both. A single encoder is trained with an invariance loss over global and local views of a clip, regularized by SIGReg, which excludes collapse with a provable guarantee. The architecture reduces to an encoder and a projector, and the objective to a single hyperparameter. This formulation admits two properties. First, the cost of pretraining is governed by the number of tokens the encoder observes; uniform random token dropping renders this number small while simultaneously improving downstream accuracy. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute, and at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Second, since no asymmetry between branches is required, the encoder can be trained with block-causal attention at no measurable accuracy cost: temporal ordering becomes a property of the encoder itself. Against a compute-matched DINOv2 trained on frames of the same videos, LeVJEPA approaches the image-pretrained encoder on appearance-centric evaluation while nearly doubling its motion-centric accuracy. These results indicate that, once its computational overhead is removed, video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.

Lukas Kuhn, Lucas Maes, Giuseppe Serra et al. · 0 citations
Preprint Aug 2026

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

This work proposes a novel one-shot unlearning approach, abandoning iterative optimization in favor of a direct, exact analytical solution, and achieves state-of-the-art forgetting-utility trade-offs on TOFU-5%, TOFU-10%, MUSE-Books, MUSE-News and WMDP, significantly reducing computational overhead without sacrificing overall model performance.

Paweł Batorski, P. Spurek, Paul Swoboda · 0 citations
Aug 2026

EP-MAE: A resource-efficient masked autoencoding framework for 3D neural representation learning.

Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.

Jian Zhu, Jiale Zhao, Cheng Lin et al. · 0 citations
Preprint Aug 2026

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

A retraining-free VLM pruning framework called PORTA is introduced that derives a task- and modality-agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature-level representation utility across modalities.

Minseok Kang, Hyunwoo J. Kim, Chanyoung Kim et al. · 0 citations
Open access Jul 2026

Adaptive test-time augmentation via KL-regularized reinforcement learning for robust visual inference.

Deep neural networks often suffer significant accuracy degradation when exposed to real-world image corruptions and distribution shifts. To overcome the limitations of fixed, input-agnostic test-time augmentation (TTA), an adaptive framework is proposed that learns per-sample transformations via reinforcement learning. Augmentation selection is cast as a Markov decision process and proximal policy optimization (PPO) agents are trained to choose sample-specific transforms under a composite reward combining classifier confidence gains with a self-consistency KL-divergence penalty on the model's own softmax outputs, thereby preserving overall belief stability. On clean CIFAR-10 (1 000 samples), the adaptive ensemble raises accuracy from 88.5% (baseline) and 87.3% (static TTA) to 90.0% (+1.5 pp). On CIFAR-10-C (15 corruptions × 5 severities; 1 000 images per condition), pooled top-1 accuracy improves from 75.7% (baseline) and 74.3% (static TTA) to 76.4% (+0.7 pp), and exceeds a TENT entropy-minimization baseline (75.9%) while operating in a strictly label-free regime that updates no model weights. Per-corruption gains are consistently positive across noise, blur, weather, and compression distortions, with the adaptive policy outperforming TENT on texture and compression corruptions where input-space transforms are most effective. These findings demonstrate that learned, per-sample augmentation policies improve robustness and reliability of deep vision models under diverse image conditions, against a strong baseline classifier.

T. Mittal, A. Dubey, Dharmender Saini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.