Skip to content

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

Sep 2026 · Neural Networks · Vol 205 Pt C, pp. 109636 · 0 citations · 54 references
Medicine

TL;DR

A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.

Abstract

Few-shot open-set recognition (FSOR) presents unique challenges due to the limited labeled samples and the presence of unknown classes during inference. While recent few-shot learning methods leverage class-level textual information to enhance feature discrimination, they often overlook image-grounded descriptions that provide more fine-grained and instance-specific semantics. Moreover, many of these approaches require prior knowledge of the class name from labeled samples during inference, which may be impractical in open-world scenarios. In this study, we propose a multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance. MGF leverages image descriptions generated by a vision-language model as the textual supervision to guide the learning of semantic features from images. A graph convolutional network is then used to fuse semantic and visual features, enabling effective intra-class information propagation and improving discrimination between known and unknown classes. We jointly optimize a contrastive and a semantic alignment loss to promote intra-class compactness and inter-class separability. Extensive experiments on several few-shot learning benchmarks demonstrate that MGF achieves superior open-set recognition and competitive closed-set classification performance.

View source

Similar papers

Preprint Sep 2026

Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these met...

Yi-Jun Hu, Heng Fan, Li-Bo Zhang · 0 citations
Open access Aug 2026

Few-shot image classification algorithm based on deep learning and feature fusion

Few-shot image classification remains difficult because a model must identify novel classes from only one or a few labeled examples while preserving discriminative local information. Metric-learning methods based on Earth Mover’s Distance (EMD) improve local correspondence by representing an image as a set of regional...

Huie Zhang, Mary Jane C. Samontet · 0 citations
Conference Aug 2026

MASF-Net: efficient linear attention guided few-shot fine-grained image recognition

Few-shot fine-grained image classification (FS-FGIC) aims to distinguish visually similar subcategories with only a handful of labeled examples, posing significant challenges due to subtle inter-class differences and large intra-class variations. Existing methods often fail to fully leverage complementary information f...

Jinyu Wang, Bing-Xin Xu, Weiguo Pan et al. · 0 citations
Open access Sep 2026

Dual Adaptive Visual-Semantic Prompt Collaboration for Generalized Zero-Shot Learning

Generalized zero-shot learning (GZSL) addresses the challenging task of recognizing both seen and unseen classes by leveraging shared semantic knowledge. A core challenge in this domain is achieving robust visual-semantic alignment to transfer knowledge from seen classes to novel classes. Current state-of-the-art metho...

Huajie Jiang, Zheng-Xian Li, Yuankai Qi et al. · 0 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations
Preprint Sep 2026

Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning

Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this...

Enrico Pallotta, Sina Raoufi, Lars Doorenbos et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.