Skip to content

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

SPMC, a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention, demonstrates that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.

Abstract

Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attention heads, we investigate whether reasoning can be improved by guiding only the heads most responsive to visual evidence grounding. We propose Selective Probability Mass Concentration (sPMC), a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention. sPMC treats normalized attention over visual tokens as a spatial probability distribution and encourages the probability mass to be assigned to semantically relevant regions using segmentation-derived spatial priors. Adaptive Head Selection restricts this guidance to visually responsive heads while leaving the remaining heads unconstrained to preserve their complementary functions. Across 6 multimodal benchmark suites, sPMC achieves an average zero-shot improvement of 3% and gains of up to 11.3% across multiple MLLMs while regularizing only 3%-15% of their attention heads. These results demonstrate that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.

View source

Similar papers

#computer vision Preprint Sep 2026

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-base...

Davide Testa, Hugh Mee Wong, Alessandro Lenci et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before l...

Tian-Hang Guo, Yu-Lin He, Wei Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models

Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes...

Rafi Ibn Sultan, Xiang-Yu Zhou, Mohammad O. S. Chowdhury et al. · 0 citations
Preprint Aug 2026

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

Heng-Yuan Xu, Wei Cheng, Yu-Meng Ji et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

Whether attention weights faithfully reflect model reasoning has been actively debated in NLP, yet this question remains largely unexplored for the visual modality in Vision-Language Models (VLMs). We address this gap through causal perturbation analysis on current VLMs, evaluating both the comprehensiveness and suffic...

Xu-Rui Song, Wei-Shi Wang, Zhong-Qi Yue et al. · 2 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.