Skip to content

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Jul 2026 · arXiv.org · Vol abs/2607.13454 · 0 citations · 65 references
Computer Science

TL;DR

GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure, and is introduced a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding.

Abstract

Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent continuous geometric information. While recent methods use latent representations to enhance reasoning, relying on a single latent type cannot adapt to the diversity of spatial tasks, leading to misalignment in complex geometric scenarios. To address these limitations, we propose GeoAnchor, an interleaved text-latent reasoning framework. GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure. These components are recombined in a structured space to construct local evidence while capturing global context, enabling dynamic and interpretable reasoning. Furthermore, we introduce a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding. Extensive experiments on diverse and complex 3D reasoning tasks demonstrate that GeoAnchor outperforms the state of the art, validating its effectiveness and generalization capabilities.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.

Jun Du, Fernando Ropero, Erkin Turkoz et al. · 0 citations
Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation.

Lucy Lin, Ayush Jain, Yifan Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning

Quantitative spatial reasoning in visual-language models (VLMs) aims to infer spatial distances and directional relationships among objects in 3D space from a 2D image and a natural language query. Despite recent progress, VLM spatial reasoning remains brittle under distribution shifts, largely due to the high cost of 3D supervision. As a result, models often produce inconsistent or contradictory predictions when faced with novel object configurations or rephrased spatial queries, revealing a misalignment between learned representations and underlying geometry. To address this, we propose TTL-SR, a geometry-aware Test-Time Learning framework for quantitative Spatial Reasoning that leverages geometric consistency constraints and unlabeled test data to adapt models to target domains. Specifically, TTL-SR augments the input query with geometrically coupled auxiliary queries, filters unreliable predictions via adaptive geometric triggering to construct structured token-level pseudo-labels, and updates model parameters under a geometry-aware multi-objective loss using only test data. Experimental results demonstrate that TTL-SR significantly boosts spatial reasoning performance, yielding 6.47% and 9.41% accuracy gains for Qwen3-VL-4B-Instruct and SpatialRGPT-VILA-1.5-8B on Q-Spatial-ScanNet dataset, respectively.

Ge-Ge Zhang, Shuai-Cheng Niu, Gang Dai et al. · 0 citations
Conference Jul 2026

3D vision-language question answering with explicit scene graphs and local topology priors

TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.

Kaixin Wu, Kunlin Zhou, Boxin Li et al. · 0 citations
Conference Aug 2026

Enhancing Spatial Understanding in Vision-Language Models via Curriculum Learning

The development of Embodied AI urgently necessitates high-fidelity environment modeling enriched with spatial context. However, existing 3D semantic scene understanding methods predominantly focus on isolated instance-level labels or high-dimensional semantic vector embeddings, lacking effective modeling of explicit spatial-semantic relationships between objects within a scene. To address this issue, we use a method that guides Vision-Language Models(VLMs) to learn scene spatial layout relationships via Supervised Fine-Tuning (SFT). First, based on the InteriorGS dataset, we construct a Visual Question Answering (VQA) dataset spanning 100 scenes, comprising RGB images and grounding images with 2D bounding box prompts. Within this dataset, we systematically annotate the spatial-semantic relationship graphs among visible instances. Second, we utilize the parameter-efficient fine-tuning strategy of Low-Rank Adaptation (LoRA) to enhance the spatial relationship reasoning capabilities of the baseline model, Qwen2.5-VL-7B-Instruct. Furthermore, we design an easy-to-hard, three-stage curriculum learning scheme: progressing from single-image single-instance relationship reasoning, advancing to single-image multi-instance relationship understanding, and ultimately achieving global spatial layout perception across continuous frames. Comprehensive evaluations demonstrate that our method enables the model to effectively comprehend and output spatial-semantic relationship triplets in a predefined format, significantly outperforming the baseline model in inter-instance spatial-semantic reasoning. Our research validates the feasibility of endowing existing VLMs with preliminary spatial intelligence via SFT, laying the foundation for constructing next-generation semantic scene representations enriched with spatial relational information.

Zhe Zhong, Qin-Yuan Ren · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.