Skip to content
Preprint

Uncertainty-Aware Art-Historical Dating with Vision-Language Models

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

This work evaluates several pretrained vision models on a temporally controlled Wikidata corpus of artworks and shows that these models contain usable temporal information, however, a qualitative analysis indicates that this temporal knowledge is shaped by various biases.

Abstract

Museum and archival datasets do not mirror historical artistic production, but materialize the contingent histories of collecting, preservation, cataloging, and digitization. This has direct consequences for interpreting pretrained image representations: they may appear to encode historical time while actually encoding the institutional conditions under which objects become visible as data. We describe this phenomenon as temporal entanglement and investigate it by formulating artwork dating as an uncertainty-aware regression task over frozen image embeddings. We evaluate several pretrained vision models on a temporally controlled Wikidata corpus of artworks. Our results show that these models contain usable temporal information, with Vision-Language Models (VLMs) outperforming purely visual self-supervised baselines. However, a qualitative analysis indicates that this temporal knowledge is shaped by various biases.

View source

Similar papers

Preprint Jul 2026

Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations

This work introduces CANVAS (Contrastive Art-aware Network for Vision-Language Alignment with Sheaves), a framework for learning relation-aware multimodal representations inspired by sheaf theory, and evaluates on three newly introduced benchmarks of artworks for multi-relational art understanding.

Ludovica Schaerf, Antonio Purificato, Piera Riccio et al. · 1 citation
Preprint Jul 2026

Using Hierarchical Controlled Vocabularies to Understand CLIP Retrieval Failures in Historical Photo Collections

It is found that visual coherence and text-image alignment are nearly uncorrelated across terms and jointly separate distinct failure modes, and fine-tuning improves retrieval overall, but its gains favour shallower terms in the hierarchy, where text-image alignment improves most, beyond what concept frequency explains.

Ratan J. Sebastian, Anett Hoppe, Christoph Rippe et al. · 0 citations
Preprint Aug 2026

MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with formal analysis, grounded historical interpretation, or affective characterization. We argue this is not only a model but also a dataset limitation. Existing art datasets are single perspective resources, where no dataset provides narrative, formal, emotional, and historical perspectives simultaneously for the same artworks. We introduce MMArt, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations. Two complementarity analyses establish that perspectives encode genuinely distinct information. A generative analysis shows that formal analysis descriptions best preserve compositional style, and historical descriptions carry strong affective signal in reconstructed images. A discriminative retrieval analysis reveals task-asymmetry: narrative descriptions drive retrieval (R@1 = 44.0%), while formal descriptions, strongest for reconstruction, are nearly nondiscriminative at retrieval scale (R@1 = 7.8%). Leave-one-out analysis further confirms that historical descriptions are the least replaceable perspective across both tasks. Together, the two analyses establish that no single perspective suffices for all tasks, directly motivating MMArt multi-perspective design. The dataset, code, and additional information are available at https://shuaiwang97.github.io/MMArt/.

Shuai Wang, Wangyuan Ding, Yixian Shen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Conducting Stylistic Analysis of Paintings through an Art-History Agent

This approach converts detailed visual features into descriptive terms, addressing a key challenge in art history, and connects the use of images as data with the semantic concerns of humanists, establishing vision-based computational art history as an area for future growth.

M. Walton, Astrid Harth · 0 citations
Preprint Aug 2026

Learning visual representations for compositional analysis of artworks and photographs

Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization. Code and pre-trained models are available on GitHub.

F. Behrad, T. Tuytelaars, Johan Wagemans · 0 citations
May 2026

Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

Across OLMo-2, Llama-3.1, and Qwen-3, under both MEMIT and AlphaEdit and in batch and sequential regimes, Moir consistently extends preservation in the most vulnerable domains, suggesting that aligning the preservation distribution with the model's operative distribution is a key factor in non-destructive editing and that the model itself may be the most accessible source of that distribution for deployed systems.

Jea Kwon, Jiwon Kim, Dong-Kyum Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.