MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs
These findings suggest MLLMs function as language-guided passive observers advocating for perceptually-independent architectures that decouple sensory agency from linguistic dominance, and Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck.