Preprint
Aug 2026
Vision-Language Grounding as Bidirectional Concept Correspondence
This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.
Jieyu Zhang, Ziqi Gao, Luke S. Zettlemoyer et al.
· 0 citations