Sep 2026· IEEE transactions on circuits and systems for video technology (Print)· 0 citations· 63 references
Computer Science
TL;DR
A high-performing teacher is built that makes navigation evidence selection explicit and compressible, and a compact student is trained by transferring both where to attend and what to do, then further match action distributions during fine-tuning.
Abstract
Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deploy due to their large parameter counts and computational requirements. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence selection explicit and compressible. The teacher introduces a small set of learnable query slots to extract global and local action-sufficient navigable evidence from panoramic observations via a Navigable Query Generator, then progressively grounds these evidence tokens to the instruction with an Instruction-Query Aligner for policy prediction. Second, using this explicit query bottleneck as a distillation interface, we train a compact student by transferring both where to attend and what to do. We distill the teacher's global and local navigable queries with a navigation-aware token-adaptive objective, then further match action distributions during fine-tuning. Experiments on standard VLN benchmarks demonstrate that our student nearly matches the teacher's navigation performance while reducing the number of parameters by 93.65% compared to the teacher.
Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future obs...
Khang Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen et al.· 0 citations
Q-CueGraph maps a question and an image representation to budgeted, coordinate-level observations for a frozen reader, and reaches 92% of full-image ANLS on InfographicVQA from about half the image area.
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural language instructions. Current zero-shot VLN-CE methods either rely on pre-trained waypoint predictors or require multiple queries to large models per step. To address prohi...
Shi-Qi Pan, Qi Zheng, Hanmeng Sun et al.· 1 citation
Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, (OOD)) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM...
Rajat Modi, P. Pathak, Xin Liang et al.· 0 citations
LiteSearch-VL is a low-compute recipe for Qwen3-VL-2B and Qwen3-VL-4B that uses only released OpenSearch-VL trajectories, parameter-efficient LoRA adapters, and synthetic step-level preferences, identifying answer verification, not search depth, as the next bottleneck for small multimodal agents.
Saeed Khaki, Nima Safaei, Kamal Ginotra· 0 citations