An image set that systematically untangles global shape, internal parts, and texture information is created, and human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives is compared, revealing systematic and persistent differences between human and machine vision.
Abstract
Deep neural networks (DNNs) are promising computational models for understanding visual object recognition. Yet, whether DNNs use similar visual cues for object recognition as humans do remains unknown. We created an image set that systematically untangles global shape, internal parts, and texture information, and compared human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives. No DNNs replicated humans’ cue-reliance profile, including those with recurrence or specialized training. Fine-tuned text-image contrastive-trained models, regardless of architecture, were most human-like overall, but lost their human-alignment when the global shape was disrupted. Strikingly, all DNNs substantially underperformed humans when the global shape cue alone was critical to object recognition. Furthermore, alignment with ventral stream neural recordings in an existing database did not predict alignment to human behavior, and model performance does not always predict its human-alignment. Together, these findings reveal systematic and persistent differences between human and machine vision.
Deep convolutional neural networks are leading models of biological vision, largely because of their strong brain alignment: their features predict neural responses better than earlier models. Yet they are believed to recognize objects differently, relying on texture where humans rely on shape and failing on perturbati...
H. Scholte, Niklas Müller, Julio Smidi et al.· bioRxiv· 0 citations
How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning...
Matteo Dunnhofer, Christian Micheloni, Kohitij Kar· 0 citations
How neural activity across the ventral visual hierarchy supports face recognition is an open question. A long-standing debate asks whether face processing, particularly in fusiform cortex, relies on face-specific computations or representations shared with broader visual recognition. Here we combine source-resolved mag...
Hamza Abdelhedi, Shahab Bakhtiari, Karim Jerbi· bioRxiv· 0 citations
Deep neural networks (DNNs) are known to produce erroneous results under real-world noisy inputs, presenting a major bottleneck to their use in applications where lives, safety, or significant resources are at stake. It has been commonly observed that humans are highly resilient to the noisy inputs that are challenging...
Bharath Anand, Sarada Krithivasan· Frontiers in Artificial Inte...· 0 citations
Abstract Motivation The inductive bias of a deep learning model influences the features it extracts from biological images, making model selection a critical scientific decision. We systematically compare representations learned from scratch without pre-training by convolutional neural networks (CNNs), vision transform...
Jacob I. Evarts, Jason Y. Cain, Po-Hao Chiu et al.· Bioinformatics Advances· 0 citations
This work introduces axis-aligned feature accentuation, which converts each model’s fitted encoding axis into graded stimulus perturbations that are predicted to parametrically control neural firing within and beyond the natural-image range.
Jacob S. Prince, Binxu Wang, Thomas Fel et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.