Skip to content

Do Vision Model See Like the Brain? A Comparison Across EEG Encoding Model

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

It is proposed that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers's self-attention and non-classification objectives, which reflect signal strength and persistence rather than distinct brain regions.

Abstract

Convolutional neural networks (CNNs) and vision transformers are both used to model the human visual system, but whether the two architectures diverge at a specific point in network depth is unclear. We compared six CNNs and two vision transformers by computing the Pearson correlation (r) between each model's predicted and measured EEG response at every layer or block, in ten participants viewing 200 natural images. For the transformer models, we also tested four token representations, from the classification (CLS) token alone to CLS combined with all patch tokens. CNNs showed strongest correspondence at the earliest layers, weakening at deeper layers, particularly later in the post-stimulus response. Transformers instead sustained strong correspondence at their deepest blocks, though not at their earliest ones. This advantage depended on token representation: pooled representations gave weaker peak correlations (r approx 0.48-0.51) than representations retaining all patch tokens (r=0.640 for CLIP-ViT-B/32, r=0.656 for DINOv2-ViT-B/14). Controlled comparisons showed architecture, not training objective, drove this effect: MoCo-v1 and ResNet-50 (matched architecture) performed nearly identically (r=0.673, 0.670), whereas CLIP-RN50 and CLIP-ViT-B/32 (matched objective) diverged until patch tokens were preserved. We propose that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers'self-attention and non-classification objectives. A spatial topography analysis showed a common occipital-dominant pattern across all models, indicating these differences reflect signal strength and persistence rather than distinct brain regions. Patch-preserving transformer representations sustain brain-predictive correspondence where CNNs collapse.

View source

Similar papers

Open access Aug 2026

Brain alignment in deep neural networks emerges early and independently of object classification

It is found that alignment with human fMRI, EEG, and macaque electrophysiology is already largely present at initialization, when networks classify at chance, and reaches a plateau within one to five epochs; thereafter it changes only modestly while classification accuracy continues to climb to 75%.

H. Scholte, Niklas Müller, Julio Smidi et al. · 0 citations
Open access Sep 2026

Task-optimized neural networks reveal distinct contributions of specialized and broader visual learning to neural representations of face familiarity

How neural activity across the ventral visual hierarchy supports face recognition is an open question. A long-standing debate asks whether face processing, particularly in fusiform cortex, relies on face-specific computations or representations shared with broader visual recognition. Here we combine source-resolved mag...

Hamza Abdelhedi, Shahab Bakhtiari, Karim Jerbi · 0 citations
Preprint Sep 2026

The Visual Target Matters: Learning across the Visual Hierarchy for Brain-to-Image Retrieval

Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented by pretrained vision models, whose internal representations vary in abstraction across depth. Existing methods usually train the neural encoder to recover a fixed final...

Ye Wang, Hao-Kun Ren, Hong Yu et al. · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.