Skip to content
Preprint

SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

SAIN is presented, a zero-shot framework that turns active dialogue into persistent navigation state and supports dialogue-to-state conversion as an effective zero-shot mechanism for long-horizon interactive instance navigation.

Abstract

Most existing vision-language navigation tasks assume that instructions are complete and unambiguous. However, real-world robots often encounter natural human instructions that are ambiguous, underspecified, or incomplete, requiring them to resolve such uncertainties through active questioning. Interactive Instance Goal Navigation (IIGN) requires an embodied agent to find the specific instance under an ambiguous category-level instruction through active dialogue. However, existing dialogue-enabled methods often consume oracle answers as transient textual context for immediate decisions, rather than persistent spatial or object-centric structured state. We present SAIN, a zero-shot framework that turns active dialogue into persistent navigation state. Instead of consuming oracle answers as one-step text hints, SAIN compiles them into target evidence, route-level corridor memory, and object-candidate labels. These states are stored in structured value, room, graph, and object memories, then consumed by a unified policy for frontier ranking and final target approach. On the VL-LN IIGN benchmark, SAIN improves SR from 20.2 to 25.4 and SPL from 13.07 to 14.17 over the strongest reported dialogue-enabled baseline, while requiring no task-specific policy training. The results support dialogue-to-state conversion as an effective zero-shot mechanism for long-horizon interactive instance navigation. Project website: https://zorattc.github.io/SAIN/

View source

Similar papers

Preprint Aug 2026

Hierarchical Fast-Slow ReAct Agent for Zero-Shot Object-Goal Navigation

This paper presents a hierarchical fast-slow agent that turns what the robot has already seen into the object of deliberation in zero-shot object-goal navigation, and reaches the highest success rate among the zero-shot methods compared here.

Zhaochen Lan, Zhi Yang, Yuxiang Fu et al. · 0 citations
Jul 2026

WCM: World-Cognition Model for Generalizable Human-Robot Interaction

The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.

Yuzhen Chen, K. Zhou · 0 citations
Preprint Jul 2026

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

ABot-AgentOS is presented, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration.

Jiayi Tian, Shiao Liu, Yuting Xu et al. · 0 citations
Aug 2026

ReasonWalker: Reasoning Iterative Vision-and-Language Navigation With Implicit Instructions.

Existing vision-and-language navigation (VLN) agents typically cannot infer users' implicit intentions. They are unable to leverage past experiences in persistent environments. In this article, we propose ReasonWalker, a novel navigation model designed to enable reasoning-based navigation using implicit instructions over time. To ensure persistent and efficient operation, ReasonWalker constructs and stores explicit scene maps, allowing it to learn scene associations for improved renavigation in subsequent episodes. To facilitate comprehension and reasoning over implicit instructions, ReasonWalker leverages a large language model (LLM) to jointly process user instructions, agent observations, and scene maps, generating semantic navigation tokens that guide action prediction. To train ReasonWalker, we propose a new hierarchical learning paradigm, where the model first learns navigation actions and then acquires scene associations for implicit instruction reasoning. Additionally, we provide a new implicit instruction benchmark to support training and evaluation of reasoning-based navigation tasks. Extensive experiments demonstrate the effectiveness and superiority of the proposed ReasonWalker. The project page with video presentations and code is at: https://wangxudongsia.github.io/ReasonWalker-Web/.

Xudong Wang, Baicheng Liu, Jiahua Dong et al. · 0 citations
Preprint Aug 2026

Active Perception for Embodied Disambiguation

Natural language provides robots with a flexible task interface, but target ambiguity in embodied environments arises not only from user intent; it can also result from missing taskrelevant physical evidence in the current observation. Existing interactive disambiguation methods primarily obtain additional information by asking the user, whereas occlusion, restricted viewpoints, unreadable text, and unobserved targets require the robot to actively change its observation. We propose an active-perception framework for embodied target disambiguation that uses active observation as the backbone for information acquisition and uses a vision-language model to decide, on the basis of accumulated visual evidence and interaction information, whether to continue observing, request clarification, or complete target selection. Active observation can both directly recover missing discriminative evidence and reveal object names, labels, and semantic attributes, thereby improving user clarification when it remains necessary. Real-robot experiments show that the framework combines physical information acquisition and userintent clarification within a unified embodied disambiguation process.

Yiwei Liu, Luwei Yang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.