Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception
It is found that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; and that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representa...