Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs
This work studies ViT attention heads and finds they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention, and proposes SHS-Index to quantify this specialization, showing that it distinguishes full-attention from chunk-window ViTs, and finds that it strongly tracks downstream benchmark performance.