Skip to content
Conference

3D vision-language question answering with explicit scene graphs and local topology priors

Jul 2026 · International Conference on Generative Artificial Intelligence and Image Processing · Vol 14292, pp. 142920B - 142920B-6 · 0 citations · 15 references
Engineering

TL;DR

TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.

Abstract

The development of 3D large multimodal models (3D-LMMs) has advanced research on 3D visual question answering. Yet most existing methods rely on implicit feature mapping, where point clouds or scene features are directly projected into the latent space of a language model, without explicitly modeling local spatial structure. In complex indoor environments, this design can lead to spatial judgments that are inconsistent with the actual physical layout during 3D visual question answering, a phenomenon referred to as spatial hallucination. To address this issue, this paper proposes TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors. The method begins by parsing a raw 3D scene into an explicit spatial-semantic scene graph and extracting instance-level representations that encode both semantic features and geometric location information. It then constructs a local physical neighborhood around the target object, serializes neighboring objects together with their distance information into structured priors, and injects them into the multimodal reasoning sequence as conditional context. Under the current experimental setting on the Replica dataset, the results show that the proposed method achieves strong performance on spatial-relation question answering while also alleviating spatial hallucination to a certain extent. These findings suggest that explicit local topological priors can improve scene consistency in 3D visual question answering.

View source

Similar papers

Open access 2026

Query-Driven Evidence Retrieval for Efficient 3-D Question Answering

The integration of Large Vision-Language Models (LVLMs) with 3D scene understanding has shown great promise. However, existing 3D Visual Question Answering (3D QA) paradigms face severe bottlenecks. Directly feeding dense, unconstrained multi-view video streams or full point clouds into LLMs incurs prohibitive computational overhead and attention dilution, rendering models highly susceptible to spatial disorientation and visual hallucinations. To address these challenges, we propose Opti3D, a training-free, plug-and-play structured evidence construction framework that retrieves compact query-relevant visual evidence for off-the-shelf Video LLMs. Specifically, Opti3D constructs a global Bird’s-Eye View (BEV) map to preserve macroscopic spatial layout and further organizes local object observations through explicit geometry-aware cross-view deduplication. Instead of treating every sampled frame as an independent visual input, Opti3D groups redundant 2D observations that correspond to the same physical instance and then selects compact, viewpoint-informative local evidence using a multi-dimensional view-quality score. In this way, complex 3D roaming videos are converted into a structured visual evidence set containing a global BEV context and non-redundant local instance views. Extensive experiments on challenging 3D QA benchmarks demonstrate that Opti3D achieves competitive performance among training-free or frozen-backbone Video LLM baselines, while remaining below specialized 3D LLMs trained with task-specific 3D-language supervision. The results show that explicit geometry-aware view deduplication reduces redundant visual inputs and provides more reliable evidence for viewpoint-ambiguous questions without requiring full supervision or task-specific training.

Huihui Liu, Haoyang Wu · 0 citations
#artificial intelligence Preprint Sep 2026

GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.

Jun Du, Fernando Ropero, Erkin Turkoz et al. · 0 citations
Aug 2026

Boosting scene captioning with pairwise semantic spatial reasoning and contextualized large language model integration

This work extends its previous work in scene graph generation to integrate image captioning, incorporating large language models to interpret scene graph structures and produce high-quality captions encompassing complex actions and relationships.

Anfel Amirat, N. Baha, Lamine Benrais · 0 citations
Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation.

Lucy Lin, Ayush Jain, Yifan Liu et al. · 0 citations
Preprint Aug 2026

Inter-3D VQA: A Roadside Multimodal Benchmark for 3D Spatiotemporally Grounded Visual Question Answering

Recent advances in visual question answering (VQA) and multimodal large language models (MLLMs) have enabled natural-language reasoning over traffic scenes. However, existing benchmarks are largely built from ego-vehicle views or 2D roadside videos, limiting their ability to evaluate 3D-grounded reasoning over real-world distances, trajectories, infrastructure topology, and safety-critical interactions. We introduce Inter-3D VQA, a large-scale roadside multimodal benchmark for 3D spatiotemporally grounded VQA at intersections. Built from synchronized point clouds and multi-view images, Inter-3D VQA contains 407K QA pairs covering lane-level positions, object relationships, motion patterns, and near-miss-oriented interaction reasoning. We further propose Inter-Geo, an MLLM baseline that integrates object- and scene-level aligned LiDAR representations, and Inter-Metrics, a unified evaluation framework for textual consistency, numerical accuracy, and semantic correctness. Experiments show that Inter-Geo outperforms image-based VLMs, especially on grounded spatial and temporal reasoning tasks. Our benchmark and codes are available at https://github.com/ASU-Suo-Lab/Inter-3D-VQA .

Shaozu Ding, Linan Song, Dajiang Suo · 0 citations
Jul 2026

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure, and is introduced a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding.

Hao Li, Han Fang, Zixin Pan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.