Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models
This work finds that the wording of the question affects VLMs in two opposite ways: verbose paraphrasing reduces drift variance by 70--81% on the 8B models and the practical recipe---pad the prompt---further yields measurable gains in accuracy, even under image corruption.