Skip to content

Author

Wei Emma Zhang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

PeaCap: Patch-Level Retrieval for Lightweight Retrieval-Augmented Image Captioning

Retrieval-augmented image captioning aims to improve caption quality by grounding generation in external evidence, but most prior systems retrieve evidence using coarse whole-image similarity, which can miss small or rare objects in cluttered scenes. We propose PeaCap, a patch-based retrieval-augmented captioning framework that explicitly studies how retrieval granularity affects the quality of retrieved object evidence and downstream caption generation. PeaCap decomposes a query image into patches, performs patch-level image-to-image retrieval to obtain object tags, and fuses the retrieved tags with the whole-image embedding via a lightweight cross-attention module and an alignment loss to robustly prompt a frozen LLM. Analyses on retrieval (encoder choice, patch-vs.-whole retrieval, and patch-grid ablations) show that patch-level retrieval can improve object coverage, and experiments on COCO and out-of-domain benchmarks demonstrate competitive captioning performance under a lightweight training setup.

Robin Viltoriano, Wei Emma Zhang, Hu Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.