Skip to content

Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

Jul 2026 · arXiv.org · Vol abs/2607.08321 · 0 citations · 44 references
Computer Science

TL;DR

It is found that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; and that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations.

Abstract

In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. However, it remains unclear to what extent this alignment persists for visual tasks that extend beyond the canonical object recognition paradigm based on well defined semantic content. In this study, we diverge from the common object-centric view by focusing on another aspect of vision: texture perception. We consider textures of different complexity generated with three different algorithms from the same source images. Using a rank-based statistic, we quantify the information encoded in the internal representations of a CNN and three Vision Transformers (ViTs), and we compare the similarity of these representations to those inferred from human psychophysics data. We find that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations. Taken together, these results suggest that ViTs may capture more faithfully than CNNs how texture patterns are visually processed by humans, and that the representations of texture stimuli in computational models may be driven by the network architecture.

View source

Similar papers

Preprint Aug 2026

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remains mechanistically unclear how they encode low-level features, given their lack of inductive biases: ViTs process information globally rather than relying on local struc...

Vaishnavi B Mohan, Vijayakrishna Naganoor, Yashas Annadani et al. · 0 citations
Review Open access 2026

From Convolution to Attention and Beyond: A Systematic Review of Modern Vision Architectures

A critical review of computer vision, illustrating how architectural design, learning paradigms, and evaluation practices have co-evolved over time to facilitate more flexible and scalable systems, and outlining new research directions.

Amitabha Chakrabarty, Azwad Aziz, Anika Tahsin et al. · 0 citations
Jul 2026

Shape Matters: Few-Shot Object Classification From High-Information Contour Features.

This work proposes a biologically inspired model that addresses both issues by classifying images based on transformation-invariant local shape key features, demonstrating strong robustness and greater capacity to generalize to unseen distributions, bringing it closer to human-like recognition capabilities.

Maria Osório, Alexandre Bernardino, Andreas Wichert · 0 citations
Open access Aug 2026

Systematic image perturbations reveal persistent gaps between human and machine vision

An image set that systematically untangles global shape, internal parts, and texture information is created, and human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives is compared, revealing systematic and persistent differences between human and machine vis...

Mugihiko Kato, Biyu J. He · 1 citation
Open access Jul 2026

The representational geometry of naturalistic textures in macaque V1 and V2

Summary Our understanding of visual cortical processing has relied primarily on studying the selectivity of individual neurons in different areas. A complementary approach is to study how the representational geometry of neuronal populations differs across areas, which can reveal encoding strategies difficult to infer...

Abhimanyu Pavuluri, Adam Kohn · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.