Attending to Multimodal Generation One Token at a Time
This work introduces multimodal tasks that require explicit switching between visual and textual context within a single response and proposes a simple test-time intervention to boost attention to the relevant modality at the right time, significantly improving multimodal task performance.