StaG-CoTMR: Paraphrase-Consistent Zero-Shot Composed Image Retrieval via Finite Edit Graphs and Consensus-Guided Ranking
Abstract
Zero-shot composed image retrieval (ZS-CIR) methods often use large vision-language models (LVLMs) to represent an image-text query comprising a reference image and a modification instruction. However, semantically equivalent instructions can alter the generated representations and final rankings. An individual retrieval route may fail: image-only evidence may not capture substantial edits, whereas text-only evidence may omit the identity and context of the reference image. We propose StaG-CoTMR, a zero-shot framework that addresses these problems through semantic normalization and multi-route retrieval. A query-local finite edit semantic graph maps each instruction into a shared, role-aligned representation of preserved content, edit operations, and target states, reducing semantic drift across equivalent instructions. Guided by this representation, image-dominant, text-dominant, and CoTM-R-composed routes retrieve complementary candidates, which are integrated through Consensus-Guided ranking. On FashionIQ, StaG-CoTMR improves average Consistent Success Rate at rank 10 from 28.59% to 38.24% over CoTMR, while average Recall@10 increases from 38.25% to 39.94%. These results support the combination of stable query semantics and complementary retrieval evidence.