Skip to content
Review Open access

Mathematical Optimization and Advanced Algorithms for Few-Shot and Zero-Shot Visual Learning: An Optimization-Centered Review

Sep 2026 · Mathematics · Vol 14, pp. 3407 · 0 citations · 38 references

TL;DR

A unified comparison of attribute-based ZSL, episodic meta-learning, metric and prototype estimators, graph and optimal-transport inference, generative any-shot models, and adaptation of vision–language models is analyzed.

Abstract

Few-shot learning (FSL) and zero-shot learning (ZSL) are usually studied as separate problems, yet both require prediction when class-specific evidence is absent or scarce. This review analyzes their shared difficulty from an optimization perspective. Instead of grouping studies only by architecture, it tracks four common coordinates: the information available to the learner, the variables estimated from that information, the objectives and constraints, and the numerical solvers. These coordinates support a unified comparison of attribute-based ZSL, episodic meta-learning, metric and prototype estimators, graph and optimal-transport inference, generative any-shot models, and adaptation of vision–language models. The synthesis exposes recurring trade-offs rather than a universally preferable family: flexible updates increase estimator variance; tractable task-time solvers inherit representation bias; query batches can improve inference while changing the protocol; and strong pretrained priors reduce target-data requirements while making the origin of task evidence harder to audit. Canonical objectives are distinguished from simplified review formulations and prospective research targets. The framework also clarifies the progression from explicit semantic mappings to local adaptation around pretrained image–text representations. Three priorities emerge: model selection without extra validation labels, safe use of uncertain pretrained knowledge, and stable parameter-efficient adaptation. Under this view, FSL and ZSL are connected structured-estimation problems rather than an inventory of unrelated algorithms.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational arc...

Neeraj Yadav · 0 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations
Preprint Aug 2026

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

This work proposes a novel one-shot unlearning approach, abandoning iterative optimization in favor of a direct, exact analytical solution, and achieves state-of-the-art forgetting-utility trade-offs on TOFU-5%, TOFU-10%, MUSE-Books, MUSE-News and WMDP, significantly reducing computational overhead without sacrificing...

Pawel Batorski, P. Spurek, Paul Swoboda · 1 citation
#artificial intelligence Preprint Sep 2026

Sample-Conditioned Representation Selection for Audio Few-Shot Learning

Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propos...

Feng-Rui Liu, Ning-Xin Shen, Yi Li et al. · 0 citations

UvA-DARE (Digital Academic Repository) Elastic ViTs from Pretrained Models without Retraining

SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retrainin...

Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort et al. · 0 citations
Preprint Sep 2026

Technical note on: Zero-Training Feature-Space Alignment via Information Geometry

Deep vision models often degrade under distribution shift. Test-time adaptation can improve robustness but typically requires iterative optimization, hyperparameter tuning, and multiple forward-backward passes. We propose Zero-Training Fisher Geometry Alignment (ZFGA), a closed-form method that improves robustness unde...

Behraj Khan, T. Syed, Syed Ahmad Chan Bukhari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.