Skip to content
Open access

Human–AI Co-Navigation for Indoor Object Search under Uncertainty

2026 · AHFE International · Vol 205 · 0 citations

TL;DR

A human–AI collaboration framework that utilizes a Vision-Language Model (VLM) as the perceptual and semantic backbone of a navigation agent and provides insights into the role of minimal human input in VLM-based assistive navigation systems is proposed.

Abstract

Assistive technologies for people with visual impairments increasingly use artificial intelligence to support object-finding and navigation in indoor environments. Yet fully autonomous perception remains unreliable in such settings, as indoor spaces are visually complex, only partially observable from the user’s current viewpoint, and subject to continuous change. Our work takes the position that effective assistive navigation is inherently collaborative; the system performs continuous perceptual processing, while the user provides occasional natural-language guidance when the search becomes uncertain or inefficient. To this end, we propose a human–AI collaboration framework that utilizes a Vision-Language Model (VLM) as the perceptual and semantic backbone of a navigation agent. A human user, modeled by a simulated intervention controller, provides sparse and structured guidance, which is integrated with the VLM to update its semantic search hypotheses toward the likely location of the target object. Evaluation is conducted in the Habitat simulator on photorealistic scenes from the Habitat-Matterport3D dataset. Experiments analyze how human guidance affects task success and navigation efficiency, showing that guidance is most effective when it corrects the VLM's misaligned semantic search hypotheses, providing insights into the role of minimal human input in VLM-based assistive navigation systems.

Read PDF

Similar papers

Communicating AI Uncertainty in Assistive Navigation for People with Visual Impairments

Visual impairments affect upwards of 2.2 billion people worldwide. As AI systems increasingly support navigation for people with visual impairments, how uncertainty is communicated becomes critical. Prior work shows that communicating uncertainty can improve trust calibration and decision-making, yet it remains underexplored in assistive navigation. This project develops an uncertainty-aware assistive navigation architecture that integrates AI uncertainty into auditory guidance during real-time scene descriptions. Rather than using explicit confidence statements, the prototype embeds uncertainty into speech via variations in tone, pacing, and emphasis. The prototype combines a real-time collision-warning module with a semantic reasoning layer powered by a large language model (LLM). When generating scene descriptions, token-level uncertainty is mapped to auditory prosodic cues, enabling users to implicitly gauge the system’s confidence without disrupting navigational task flow. This work presents a high-fidelity prototype that treats AI confidence as an interaction design feature, illustrating how model uncertainty can be rendered perceptible in assistive navigation and reframed as a human-factors design parameter.

Hayden Shaffer, Aahil Shaikh, He Zhang et al. · 0 citations
Jul 2026

Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments

Foundation-model-based vision-language navigation (VLN) has advanced autonomous robot navigation by enabling robots to interpret natural-language instructions, identify semantic goals, and follow user-specified behavioral rules. However, existing VLN systems rely heavily on cloud-hosted foundation models for language understanding and semantic grounding, limiting their applicability where network connectivity is unavailable and reliable metric goal localization is required. Although recent small language models (SLMs) enable fully onboard inference, their suitability for navigation instruction decomposition has not been systematically evaluated. This paper makes three contributions toward fully onboard VLN for outdoor environments. First, we present the first systematic benchmark of 17 edge-deployable SLMs against 4 online APIs for robotic navigation instruction decomposition, evaluating accuracy and latency on human-annotated instructions across three computing platforms and providing practical guidance for selecting onboard language models. Second, we propose a lightweight hybrid semantic-geometric goal localization framework that combines open-vocabulary object detection, prompted segmentation, and LiDAR geometry to estimate metric goals, while maintaining visual bearing guidance when reliable geometric observations are unavailable. Third, we integrate these advances into Edge-BehAV, a fully onboard extension of the BehAV architecture that enables cloud-independent behavior-guided navigation. Experimental results show that the best offline SLM matches the instruction decomposition performance of the strongest cloud API while running approximately 9x faster and without network connectivity. The proposed goal localization framework reduces mean goal-distance error from 2.05 m to 0.20 m at lower computational cost, and the complete system succeeds in 31 of 32 closed-loop outdoor trials.

Ali Salmasi, Xian-Jia Yu, Tomi Westerlund · 0 citations
Jul 2026

VoLN: Vision-Only Long-Horizon Navigation - Paradigm, Benchmark, and Method

This work instantiates VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark that combines long-horizon goal-directed flight, continuous 3D motion, large viewpoint changes, and context-dependent beacon selection and reveals substantial remaining challenges in long-horizon evidence integration, cross-view goal matching, and closed-loop stability.

Jiabin Lou, Haopeng Wang, Yuanshuai Wang et al. · 0 citations
Preprint Aug 2026

Assistant Placement Aria: A Benchmark for Egocentric Placement Assistance

Human assistance in robotics spans around several tasks such as navigation, object manipulation, and placement, where a key challenge is selecting target destinations that align with human intentions or preferences. We focus on this challenge in the context of Virtual Placement (VP), the task of identifying all plausible target locations given scene context and human-centric constraints. This differs from traditional placement tasks that typically focus on a single, predefined target location. The VP problem is complex, as it requires both global and local reasoning about the scene's geometry, semantics, and plausibility. To address this gap, we introduce {\bf Assistant Placement Aria}, the first benchmark to explore diverse aspects of VP, including global, local, and human-centric constraints. It contains both synthetic and real indoor scenes annotated for three tasks: (i)~2D Panel Placement, (ii)~Sitting Suggestion, and (iii)~TV Placement. Each scene includes 2D images, a 3D point cloud, and a textual description of the objects within the scene. By contributing this benchmark, we aim to encourage further research in this underexplored and challenging field that is critically dependent on relevant data. We also evaluate several foundation models for object detection and segmentation on our benchmark.

Amir Belder, Goncalo Dias Pais, R. Vivanti et al. · 1 citation
Conference Jul 2026

Vision-based Navigation Models for Autonomous Rovers: Experimental Analysis and Comparison

Traditional navigation algorithms rely heavily on the fusion of heterogeneous data from different sensor modalities for precise localisation and navigation tasks. However, these approaches lack generalisation and interaction with human operators, particularly in navigating unknown environments and human language-based task specifications. In contrast, the foundation models trained on internet-scale data proved higher generalisation capabilities and show an emerging trend in zero-shot learning for navigation tasks. Furthermore, these generic models can close the perception-planning loop through common sense reasoning, applicable to both language-based tasks and open-vocabulary visual recognition. However, due to the scarcity of robotics data for these vision-language-action-based approaches and the lack of clear evaluation protocols for these models, including safety guarantees, it is hard to evaluate the real benefits of these models. In this work, we analyse three different image-based navigation models: ViNT, NoMaD, and GNM. We propose specific metrics to quantitatively evaluate the visual-action-based navigation methods in indoor, outdoor, and simulation environments. In each scenario, these methods are adapted to a different robotic setup than the one provided in the original papers, providing an opportunity to benchmark the generalizability of these methods. The test setups and the associated codebase are available on the project webpage 1.

Szilárd Molnár, Levente Tamás · 0 citations
Review Open access 2026

A Systematic Review of Agentic AI for Autonomous Navigation: SLAM-Based Intelligent Agents

: Autonomous navigation poses a key challenge in Artificial Intelligence (AI), necessitating agents to plan and execute actions in complex, partially visible surroundings. Simultaneous Localization and Mapping (SLAM) facilitates autonomous navigation of robots and vehicle objects to construct an unfamiliar environment map while concurrently monitoring their inside position. This systematic review investigates the nascent convergence of agentic AI, defined by goal-oriented autonomy, with adaptive decision-making and reasoning, with SLAM-based navigation systems. This paper utilized Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) methodology, which concentrated on peer-reviewed articles published in (2017–2026), particularly in SLAM-based intelligent agents utilized for autonomous navigation. SLAM has transformed from a geometry-based localization framework into an advanced perceptual and reasoning paradigm for autonomous navigation. Recent advancements in agentic AI, semantic perception, multimodal learning, and embodied foundation models have facilitated autonomous agents in progressing from passive mapping to context-aware decision-making and goal-directed navigation. The emergence of agentic AI and embodied AI has revolutionized SLAM into a spatial world model that facilitates perception, memory, reasoning, and autonomous decision-making. Contemporary research emphasizes lifelong SLAM, collaborative multi-agent mapping, semantic world modelling, and the integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) for intelligent autonomous agents. Consequently, SLAM has evolved from a localization instrument to an extensive cognitive framework facilitating advanced autonomous navigation systems.

Mukesh Dalal, Anterpreet Kaur Bedi, Payal Mittal · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.