Skip to content
Preprint

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

Aug 2026 · 0 citations · 32 references
Computer Science

TL;DR

A cross-modal grounding study shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.

Abstract

Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clich\'es or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clich\'es into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clich\'ed. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-B...

Chuer Chen, Zi-Chen Wang, Yi He et al. · 0 citations
Preprint Aug 2026

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textual...

Han Chen, Ming Li, Hong Jiao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation

Imag-Eval is introduced, a controlled benchmark designed to assess how T2I models ground compositional natural-language instructions into visual outputs and suggests that, for structured skills, compositional difficulty is primarily governed by the number of grounded rules and their binding to instances, rather than by...

I. Serouis, David Jaramillo Duque · 0 citations

A Typography Benchmark for Co-Creative Graphic-Design Agents

A typography-focused benchmark of 12 tasks grounded in a layered-composition corpus of 989 real-world design templates and 2,568 text elements is introduced to quantify how well future design partners can perceive, reproduce, and manipulate the ty-pographic layer designers care about most.

Jaejung Seol, Elad Hirsch, Hao-Nan Zhu et al. · 0 citations
Preprint Sep 2026

GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solv...

Chang-Peng Zhao, Yi-Ren Song, Jin-Peng Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.