Preprint
Aug 2026
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models
The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference, and that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.
Michelle Lin
· 0 citations