Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information
This work introduces Visual Retrieval Heads (VRHs), a small subset of attention heads that are causally responsible for grounding text descriptions to image regions, and shows that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads.
Chanho Park, Daehyeon Choi, Jih-Yun Lee et al.
· 0 citations