Skip to content

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

Sep 2026 · IEEE transactions on circuits and systems for video technology (Print) · 0 citations · 63 references
Computer Science

TL;DR

A high-performing teacher is built that makes navigation evidence selection explicit and compressible, and a compact student is trained by transferring both where to attend and what to do, then further match action distributions during fine-tuning.

Abstract

Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deploy due to their large parameter counts and computational requirements. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence selection explicit and compressible. The teacher introduces a small set of learnable query slots to extract global and local action-sufficient navigable evidence from panoramic observations via a Navigable Query Generator, then progressively grounds these evidence tokens to the instruction with an Instruction-Query Aligner for policy prediction. Second, using this explicit query bottleneck as a distillation interface, we train a compact student by transferring both where to attend and what to do. We distill the teacher's global and local navigable queries with a navigation-aware token-adaptive objective, then further match action distributions during fine-tuning. Experiments on standard VLN benchmarks demonstrate that our student nearly matches the teacher's navigation performance while reducing the number of parameters by 93.65% compared to the teacher.

Read PDF

Similar papers

Preprint Sep 2026

FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation

Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future obs...

Khang Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural language instructions. Current zero-shot VLN-CE methods either rely on pre-trained waypoint predictors or require multiple queries to large models per step. To address prohi...

Shi-Qi Pan, Qi Zheng, Hanmeng Sun et al. · 1 citation
#small language model Preprint Sep 2026

Position Aware Layer Queries for Test Time Training in Vision Language Models

Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM...

Rajat Modi, P. Pathak, Xin Liang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO

LiteSearch-VL is a low-compute recipe for Qwen3-VL-2B and Qwen3-VL-4B that uses only released OpenSearch-VL trajectories, parameter-efficient LoRA adapters, and synthetic step-level preferences, identifying answer verification, not search depth, as the next bottleneck for small multimodal agents.

Saeed Khaki, Nima Safaei, Kamal Ginotra · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.