Jul 2026
Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning
This paper proposes SegAnswer, which shifts the unit of zoom-in from the popular bounding box to pixel-level segmentation mask, and demonstrates its capability for reliable pixel grounding.
Yake Wei, Yuan Wang, Fengyun Rao et al.
· arXiv.org · 0 citations