Skip to content
Preprint

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Aug 2026 · 2 citations · 70 references
Computer Science

TL;DR

WorldExam is introduced, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity, which supports unified evaluation of camera-, action-, and language-driven model paradigms.

Abstract

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.

View source

Similar papers

Preprint Oct 2026

ROWBench: Do Video Models Render What the Program Specifies?

Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instructio...

Zheng-Hui Huang, Gui-Xu Lin, Yu-Ju Tsai et al. · 0 citations
Preprint Sep 2026

World in World: Explore the World with World Models

Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded event, place observed content in the requested view, plausibly complete newly exp...

Chen-Xi Song, Yan-Ming Yang, Chi Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models

Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded bench...

Hao Wang, Tao Yu, Liu-Zhou Zhang et al. · 0 citations
Preprint Sep 2026

Programmable World Model

Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visu...

Zheng-Hui Huang, Gui-Xu Lin, Jia-Cheng Lin et al. · 5 citations · ⚡1
#artificial intelligence Preprint Aug 2026

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions...

Yuandong Pu, Le Zhuo, Sayak Paul et al. · 0 citations
Preprint Sep 2026

HappyWorld-Bench

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them...

Zhi-Qi Bai, Ju-Nai Cai, Yi-Xin Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.