Preprint
Jul 2026
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Multi-Token Localized Attention (MTLA) is proposed: a training-free, post-hoc score that measures how strongly a prediction's tokens attend to the region they claim and nearly doubles the zero-shot COCO detection AP of an open-source 8B generalist (from 20.4 to 37.0).
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen et al.
· 0 citations