Skip to content

SoftNav: Injecting 3D Scene Tokens into VLMs for Embodied Navigation

Jul 2026 · arXiv.org · Vol abs/2607.14586 · 0 citations · 37 references
Computer Science

TL;DR

SoftNav is introduced, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector, enabling transferable navigation with minimal training.

Abstract

In goal-directed embodied navigation, where an agent must locate a specified target in an unseen environment, 3D scene understanding and navigation reasoning must work in concert. Current approaches transmit 3D scene information to vision-language models (VLMs) through text, suggesting a representation gap in our tested configurations; a controlled ablation confirms that direct embedding-level transfer significantly outperforms the evaluated text serialization formats. We introduce SoftNav, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector. With the 3D encoder and VLM frozen, only ~1,200 samples and ~17M trainable parameters are needed. On HM3D-OVON, SoftNav achieves 74.2%/68.3%/66.7% SR across three splits, surpassing all prior methods in both SR and SPL; the same navigation policy transfers zero-shot to GOAT-Bench (67.2% SR), SG3D (47.2% s-SR), and real-world robot deployment without retraining or architectural modification. Injecting 3D scene tokens directly into VLMs bridges the representation gap, enabling transferable navigation with minimal training.

View source

Similar papers

Preprint Aug 2026

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

This work proposes TAMP-Nav, a unified framework for efficient embodied navigation that dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhanci...

Hongyan Feng, Sun-Lai Chen, Xuan-Yu Liu et al. · 1 citation
Preprint Aug 2026

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.

Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al. · 1 citation
Preprint Aug 2026

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal grounding framework that in...

Huosen Ou, Dong-Ni Song, Yuncong Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Shao-An Wang, Ao-Cheng Luo, Fei Huang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.