Skip to content
Preprint

PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

This paper designs a target-focused perception mechanism that identifies and enhances visual tokens of the search target, consisting of a target Enhancer and a learning strategy based on on-policy distillation and attention guidance loss, enabling the model to focus on search-target regions.

Abstract

Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial because real-world query images contain diverse information, including the search target, useful auxiliary evidence, and irrelevant visual content. This requires the model to identify and focus on the search target while selectively utilizing auxiliary evidence. In this paper, we propose \textbf{PailitaoGR}, a \emph{Latent Think-with-Images} method for generative image retrieval, which internalizes target-focused perception and selective auxiliary-evidence utilization into a the generative retrieval model, enabling \textit{Zooming without Cropping} and \textit{Reading without OCR}. Specifically, we design a target-focused perception mechanism that identifies and enhances visual tokens of the search target, consisting of a target Enhancer and a learning strategy based on on-policy distillation and attention guidance loss, enabling the model to focus on search-target regions. We also design a selective auxiliary-evidence utilization mechanism that identifies and enhances visual tokens of auxiliary evidence, including an auxiliary enhancer and an in-capacity incremental contrastive distillation strategy, enabling the model to exploit auxiliary evidence. We construct training and validation sets sampled from real-world online image-search logs. Experiments show that our method outperforms existing baselines by an average of 13.8\%, validating its effectiveness.

View source

Similar papers

Preprint Aug 2026

CoCo-IR: Contextual Composed Image Retrieval

A new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR is proposed, which interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns.

Shengcao Cao, T. Dabral, Z. Ding et al. · 0 citations
Open access Jul 2026

A Unified Multimodal Search Framework Using Generative AI and Image Understanding for Enhanced Information Retrieval

A unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.

Saeed Alzahrani, Farah Mohammad, Nazar Hussain · 0 citations
Conference 2026

Hierarchical Prompt for Task-Adaptive Composed Image Retrieval

The Task-Adaptive Hier-archical Prompt (TAHP) framework is proposed, which guides feature extraction through dynamically generated, task-specific prompts structured at three hierarchical levels: task-type, task-content, and general prompts.

Zeli Yan · 0 citations
Preprint Aug 2026

Rethinking Text-Based Image Retrieval in Specific Domain

The Semantic-Aware Fine-Tuning (SAFT) framework is proposed to address semantic compression in specific domains, which incorporates Semantic-Aware Soft-Label Supervision and Intra-modal Structural Distillation to establish a promising paradigm for domain-specific TBIR tasks.

Jingyang Tan, Shengan Yang, Yuanpeng Chen et al. · 0 citations
Preprint Aug 2026

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Real-world image search queries are multimodal and compositional: ``find this shirt in pink''specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.

Enjun Du, Siyi Liu, Zirong Chen et al. · 1 citation
Book Open access Jul 2026

PeaCap: Patch-Level Retrieval for Lightweight Retrieval-Augmented Image Captioning

Retrieval-augmented image captioning aims to improve caption quality by grounding generation in external evidence, but most prior systems retrieve evidence using coarse whole-image similarity, which can miss small or rare objects in cluttered scenes. We propose PeaCap, a patch-based retrieval-augmented captioning framework that explicitly studies how retrieval granularity affects the quality of retrieved object evidence and downstream caption generation. PeaCap decomposes a query image into patches, performs patch-level image-to-image retrieval to obtain object tags, and fuses the retrieved tags with the whole-image embedding via a lightweight cross-attention module and an alignment loss to robustly prompt a frozen LLM. Analyses on retrieval (encoder choice, patch-vs.-whole retrieval, and patch-grid ablations) show that patch-level retrieval can improve object coverage, and experiments on COCO and out-of-domain benchmarks demonstrate competitive captioning performance under a lightweight training setup.

Robin Viltoriano, Wei Emma Zhang, Hu Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.