Skip to content

Dual Retrieval Queries Fine-Tuning for Composed Image Retrieval.

Aug 2026 · IEEE Transactions on Image Processing · Vol PP · 0 citations
Medicine

TL;DR

This work proposes an asymmetric fusion mechanism to generate dual retrieval queries of different granularity, enabling the model to fully use multi-modal information and introduces a bi-directional training paradigm to ensure retrieval consistency and further exploit the triplets.

Abstract

Composed Image Retrieval (CIR) is a popular multi-modal retrieval task that aims to retrieve a target image based on a query composed of a reference image and modification text. The challenge lies in how to effectively retrieve a target image that preserves the visual content of the reference image while incorporating the changes described by the modification text. Existing CIR methods primarily employ a fusion-based strategy or a textual-inversion strategy during training. Although these methods have achieved promising results, they are limited in fully leveraging multi-modal information. This results in modality redundancy, where the retrieval process is dominated by one modality while ignoring the other. To address this issue, we propose an asymmetric fusion mechanism to generate dual retrieval queries of different granularity, enabling the model to fully use multi-modal information. Specifically, we propose a novel method termed Dual Retrieval Queries Fine-Tuning for Composed Image Retrieval (DRQ-CIR), which consists of two key components: 1) a Bilateral Multi-Modal Fusion (BMMF) module based on pre-trained VLMs, which combines the reference image and modification text to generate an enriched retrieval query; and 2) a Dual Retrieval Queries Fine-Tuning (DRQ-FT) module, which employs latent prompts to generate an enhanced retrieval query. Dual retrieval queries are used for contrastive learning with the target image to fine-tune the model and improve retrieval performance. Additionally, we introduce a bi-directional training paradigm to ensure retrieval consistency and further exploit the triplets. Extensive experiments validate the effectiveness of our proposed method on four established benchmark datasets. (Code will be available at: https://github.com/Crystal-twy998/DRQ-CIR).

View source

Similar papers

Aug 2026

Training-Free Pseudo-Fusion for Composed Image Retrieval via Diffusion Models and Multimodal Large Language Models

Composed Image Retrieval (CIR) is an emerging paradigm in content-based image retrieval that enables users to formulate compositional queries by combining a reference image with an auxiliary modality, usually text-based. This approach supports fine-grained search where the target image shares structural elements with the user-provided image while incorporating the modifications specified by the auxiliary text. Conventional CIR methods rely on multimodal fusion to combine visual and textual features into a joint query embedding, which requires training modules that align composed queries with the targets. In this work, we propose PeFuse (for pseudo-fusion), a training-free framework that leverages pretrained Diffusion Models and Multimodal Large Language Models to bridge modalities via generative conversion. We introduce two novel strategies: uni-directional and bi-directional conversion, which convert CIR into four single-modality retrieval problems. These methods reformulate CIR as either intra-modal or cross-modal single-query retrieval tasks, bypassing the need for dedicated task-specific training. Extensive experiments on standard benchmarks demonstrate that converting CIR into text-to-image retrieval tasks is more effective than alternative conversion strategies, achieving competitive or superior performance compared with state-of-the-art methods, while maintaining high flexibility thanks to replaceable components of the conversion pipeline. These results highlight the effectiveness of the pseudo-fusion paradigm for zero-shot CIR. Our code is publicly available at: https://github.com/StevenXuf/PeFuse4CIR.

Fan Xu, Luis A. Leiva · 0 citations
Conference 2026

Hierarchical Prompt for Task-Adaptive Composed Image Retrieval

The Task-Adaptive Hier-archical Prompt (TAHP) framework is proposed, which guides feature extraction through dynamically generated, task-specific prompts structured at three hierarchical levels: task-type, task-content, and general prompts.

Zeli Yan · 0 citations

CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval in Image Archives

The novel task of cross-modal query suggestion is introduced, which interactively guides users by suggesting textual refinements based on visual clusters identified in the retrieval results, and the creation of CroQS, a benchmark dataset comprising 50 diverse queries and 295 semantic clusters in generic domain.

Giacomo Pacini, Nicola Messina, Nicola Tonellotto et al. · 0 citations
Preprint Aug 2026

Rethinking Text-Based Image Retrieval in Specific Domain

The Semantic-Aware Fine-Tuning (SAFT) framework is proposed to address semantic compression in specific domains, which incorporates Semantic-Aware Soft-Label Supervision and Intra-modal Structural Distillation to establish a promising paradigm for domain-specific TBIR tasks.

Jingyang Tan, Shengan Yang, Yuanpeng Chen et al. · 0 citations
Preprint Aug 2026

CoCo-IR: Contextual Composed Image Retrieval

A new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR is proposed, which interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns.

Shengcao Cao, T. Dabral, Z. Ding et al. · 0 citations
Jul 2026

Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval

This paper introduces a Vision-Free CIR framework that addresses this challenge through two key techniques: Attribute-Augmented Hybrid Scoring, which compensates for lost visual details via explicit attribute matching, and LLM-Based Reranking, which verifies semantic consistency of top candidates.

Ryotaro Shimada, Yu-Chieh Lin, Yuji Nozawa et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.