Jul 2026
Visual Access Boundaries in Vision-Language Model Reasoning
A symbolic-attribute oracle shows that CoT can improve counting once ground-truth attributes are supplied as text, while a single-object probe-vs-decode check shows that hard attributes can be linearly recoverable from hidden states yet difficult for the model itself to output.
Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi et al.
· arXiv.org · 1 citation