Jul 2026
Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning
It is observed that the model's internal signals can indicate whether a visual step will benefit reasoning before the entire visual generation is completed, and AdaViG, a training-free adaptive visual gating method for unified multimodal reasoning, is proposed.
Wen Gao, Guanxi Lu, Di-Di Zhu et al.
· arXiv.org · 0 citations