It is found that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; and that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations.
Abstract
In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. However, it remains unclear to what extent this alignment persists for visual tasks that extend beyond the canonical object recognition paradigm based on well defined semantic content. In this study, we diverge from the common object-centric view by focusing on another aspect of vision: texture perception. We consider textures of different complexity generated with three different algorithms from the same source images. Using a rank-based statistic, we quantify the information encoded in the internal representations of a CNN and three Vision Transformers (ViTs), and we compare the similarity of these representations to those inferred from human psychophysics data. We find that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations. Taken together, these results suggest that ViTs may capture more faithfully than CNNs how texture patterns are visually processed by humans, and that the representations of texture stimuli in computational models may be driven by the network architecture.
Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remains mechanistically unclear how they encode low-level features, given their lack of inductive biases: ViTs process information globally rather than relying on local struc...
Vaishnavi B Mohan, Vijayakrishna Naganoor, Yashas Annadani et al.· 0 citations
A critical review of computer vision, illustrating how architectural design, learning paradigms, and evaluation practices have co-evolved over time to facilitate more flexible and scalable systems, and outlining new research directions.
This work proposes a biologically inspired model that addresses both issues by classifying images based on transformation-invariant local shape key features, demonstrating strong robustness and greater capacity to generalize to unseen distributions, bringing it closer to human-like recognition capabilities.
Maria Osório, Alexandre Bernardino, Andreas Wichert· Neural Computation· 0 citations
An image set that systematically untangles global shape, internal parts, and texture information is created, and human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives is compared, revealing systematic and persistent differences between human and machine vis...
Summary Our understanding of visual cortical processing has relied primarily on studying the selectivity of individual neurons in different areas. A complementary approach is to study how the representational geometry of neuronal populations differs across areas, which can reveal encoding strategies difficult to infer...
Abhimanyu Pavuluri, Adam Kohn· iScience· 0 citations