Skip to content
Conference

StaG-CoTMR: Paraphrase-Consistent Zero-Shot Composed Image Retrieval via Finite Edit Graphs and Consensus-Guided Ranking

Aug 2026 · 2026 International Conference on Computer Perception and Neural Networks (CPNN) · pp. 43-50 · 0 citations · 26 references

Abstract

Zero-shot composed image retrieval (ZS-CIR) methods often use large vision-language models (LVLMs) to represent an image-text query comprising a reference image and a modification instruction. However, semantically equivalent instructions can alter the generated representations and final rankings. An individual retrieval route may fail: image-only evidence may not capture substantial edits, whereas text-only evidence may omit the identity and context of the reference image. We propose StaG-CoTMR, a zero-shot framework that addresses these problems through semantic normalization and multi-route retrieval. A query-local finite edit semantic graph maps each instruction into a shared, role-aligned representation of preserved content, edit operations, and target states, reducing semantic drift across equivalent instructions. Guided by this representation, image-dominant, text-dominant, and CoTM-R-composed routes retrieve complementary candidates, which are integrated through Consensus-Guided ranking. On FashionIQ, StaG-CoTMR improves average Consistent Success Rate at rank 10 from 28.59% to 38.24% over CoTMR, while average Recall@10 increases from 38.25% to 39.94%. These results support the combination of stable query semantics and complementary retrieval evidence.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.