A unified comparison of attribute-based ZSL, episodic meta-learning, metric and prototype estimators, graph and optimal-transport inference, generative any-shot models, and adaptation of vision–language models is analyzed.
Abstract
Few-shot learning (FSL) and zero-shot learning (ZSL) are usually studied as separate problems, yet both require prediction when class-specific evidence is absent or scarce. This review analyzes their shared difficulty from an optimization perspective. Instead of grouping studies only by architecture, it tracks four common coordinates: the information available to the learner, the variables estimated from that information, the objectives and constraints, and the numerical solvers. These coordinates support a unified comparison of attribute-based ZSL, episodic meta-learning, metric and prototype estimators, graph and optimal-transport inference, generative any-shot models, and adaptation of vision–language models. The synthesis exposes recurring trade-offs rather than a universally preferable family: flexible updates increase estimator variance; tractable task-time solvers inherit representation bias; query batches can improve inference while changing the protocol; and strong pretrained priors reduce target-data requirements while making the origin of task evidence harder to audit. Canonical objectives are distinguished from simplified review formulations and prospective research targets. The framework also clarifies the progression from explicit semantic mappings to local adaptation around pretrained image–text representations. Three priorities emerge: model selection without extra validation labels, safe use of uncertain pretrained knowledge, and stable parameter-efficient adaptation. Under this view, FSL and ZSL are connected structured-estimation problems rather than an inventory of unrelated algorithms.
Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational arc...
G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.
Zehua Hao, Fang Liu, Qinliang Wang et al.· 0 citations
This work proposes a novel one-shot unlearning approach, abandoning iterative optimization in favor of a direct, exact analytical solution, and achieves state-of-the-art forgetting-utility trade-offs on TOFU-5%, TOFU-10%, MUSE-Books, MUSE-News and WMDP, significantly reducing computational overhead without sacrificing...
Pawel Batorski, P. Spurek, Paul Swoboda· 1 citation
Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propos...
Feng-Rui Liu, Ning-Xin Shen, Yi Li et al.· 0 citations
SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retrainin...
Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort et al.· 0 citations
Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness unde...
Behraj Khan, T. Syed, Syed Ahmad Chan Bukhari· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.