Skip to content
Book Open access

BSR-Bench: Benchmarking MLLMs’ Basic Spatial Reasoning Capabilities

Jul 2026 · Creativity & Cognition · pp. 120-129 · 0 citations · 19 references
Computer Science

TL;DR

BSR-Bench, a 2D and 3D spatial reasoning benchmark designed to probe the extent to which contemporary MLLMs can support core spatial cognitive processes, suggests that, despite recent advances, current MLLMs remain more adept at logical or linguistic abstraction than at the spatially grounded reasoning that supports human cognition and creative interaction with the physical world.

Abstract

In this paper, we present BSR-Bench, a 2D and 3D spatial reasoning benchmark designed to probe the extent to which contemporary MLLMs can support core spatial cognitive processes. In the first part of the research, we examine whether AI models can support core spatial cognitive processes and creative interaction with the physical world, using origami as a test case for spatial cognition. We then create BSR-Bench which evaluates the spatial reasoning capabilities of three widely used commercial-level MLLMs across four fundamental domains: navigation, composition, relationships, and transformation. These four areas encompass 13 prompts and 18 benchmarking metrics addressing spatial reasoning tasks, including block decomposition, orientation, and transformation, and are evaluated based on (partial or full) accuracy and consistency. These findings suggest that, despite recent advances, current MLLMs remain more adept at logical or linguistic abstraction than at the spatially grounded reasoning that supports human cognition and creative interaction with the physical world.

Read PDF

Similar papers

Preprint Aug 2026

Disentangling 3D Modeling from Spatial Reasoning

This work proposes the Disentangled Spatial Reasoner (DiSR), a simple yet effective framework that reconstructs the physical world into structured 3D evidence using off-the-shelf expert perception models and fine-tunes an LLM with LoRA to perform reasoning solely over this explicit geometric evidence.

Haoze Sun, Jie-Quan Cui, Qingshan Xu et al. · 0 citations
Jul 2026

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

It is demonstrated that training models with chain-of-thought supervision over the authors' hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.

Patrick Rim, Tom Long, Ekta Prashnani et al. · 0 citations
Preprint Aug 2026

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image, allowing errors to propagate through the reasoning chain and affect the final answer. Existing methods mainly improve spatial reasoning through training or additional spatial information, without considering whether the reasoning process itself is faithful to the model input. Our study shows that unfaithful reasoning chains significantly reduce final-answer accuracy. To address this issue, we propose a modular and training-free framework for spatial reasoning verification and correction. The framework constructs a Spatial Evidence Graph (SEG), which associates atomic spatial evidence extracted from Chain-of-Thought reasoning with visual entities, spatial relations, source steps, and visual evidence. Spatial Evidence Reliability Assessment (SERA) evaluates the reliability of visual evidence based on object existence, localization, and geometric measurements. The framework then identifies the earliest spatial evidence unit contradicted by reliable visual evidence and guides the original MLLM to revise the subsequent reasoning and final answer. Across 15 model-dataset settings, our method achieves an average accuracy of 68.94%, outperforming the compared baselines by 8.55 percentage points on average. Our code will be open-sourced.

Yang Yang, Jiawei Chen, Tairan Chen et al. · 0 citations
Preprint Aug 2026

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

This work introduces Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains, and establishes a rubric-based evaluation protocol, showing that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.

Diandian Zhang, Tingyu Song, Linbo Fu et al. · 0 citations
Jul 2026

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

The ViSTR-Bench is introduced, a novel evaluation suite designed to systematically assess whether MLLMs can perform qualitative reasoning from continuous visual cues in dynamic scenes and establishes a comprehensive four-dimensional evaluations covering Motion Perception, Spatial Relations, Outcome Prediction, and Physical Dynamics.

Han Li, Si Liu, Zehao Huang et al. · 1 citation
Jul 2026

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

ByDeWay-V2 is proposed, which integrates explicit spatial relational context alongside depth cues, expressed as human-readable predicates that serve as auditable evidence for downstream decision support, showing the framework's suitability for resource-constrained, real-time decision-support settings.

Piyush Jain, Kousik Dasgupta, Rajarshi Roy et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.