Skip to content

FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception

Sep 2026 · 0 citations · 24 references
Engineering Computer Science

TL;DR

Results show that lightweight channel-aware adaptation can recover a substantial fraction of the robustness of a much larger communication interface while broadening quality-valid operation under constrained wireless conditions.

Abstract

Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard cost. We present FreshLatent, a lightweight channel-aware latent adapter that trains a power-normalized encoder-decoder through wireless corruption while keeping the surrounding VLM frozen. We formulate deployment around a mission-conditioned perception requirement and embedded interface cost, linking channel quality and communication budget to the operating conditions under which perception remains usable. At 0 dB and the tightest communication budget, FreshLatent improves gIoU and cIoU over clean split compression by 20.79 and 20.87 points, respectively. At the most adverse evaluated SNR (0 dB), across all three communication budgets, FreshLatent recovers 63.5-69.1% of the gIoU improvement achieved by a much heavier, range-trained feature-JSCC codec. On an NVIDIA Jetson AGX Xavier in 10-W mode, FreshLatent uses 37-40x fewer encoder parameters, 7.7-9.9x lower edge-interface latency, and 8.8-10.0x lower edge-interface energy than the heavier codec. Together, these results show that lightweight channel-aware adaptation can recover a substantial fraction of the robustness of a much larger communication interface while broadening quality-valid operation under constrained wireless conditions.

View source

Similar papers

Preprint Sep 2026

Dense Coverage, Sparse Refinement: Byte-Constrained Cooperative Perception

Collaborative perception improves autonomous perception by sharing intermediate Bird's-Eye-View (BEV) features across connected agents, but dense feature exchange is difficult to deploy under strict Vehicle-to-Everything (V2X) bandwidth limits. Existing efficient methods typically either compress the full feature map u...

Melih Yazgan, Tim Muller, J. Zöllner · 0 citations
Conference Open access Sep 2026

SeGO: Sensitivity-Aware Golden Optimization for Large-Scale VLM Quantization

A cross-modal structural sensitivity asymmetry in VLMs is revealed and SeGO is proposed, a unified structural sensitivity-aware sparse optimization framework that achieves the balance among model parameter amount, quantization accuracy and scaling factors’ search efficiency on InternVL2 and LLaVA series.

Tian-Qi Zhao, Xin-Rui Cheng, Yang Su et al. · 0 citations
Preprint Sep 2026

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...

Rithvik Jonna, Man Namgung, Aakash Gurram et al. · 0 citations
Conference Aug 2026

Distilled Vision-Language Model Semantics for Edge-Deployable Content-Aware Adaptive Video Compression

With the proliferation of edge video services, content-aware adaptive video compression has become increasingly critical to balancing bandwidth efficiency and visual quality. However, conventional codec control methods mainly rely on low-level signal statistics, limiting their adaptability to high-level content semanti...

Qian Wei, En-Fang Cui, Zhi-Yuan Liang et al. · 0 citations
Preprint Aug 2026

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver. However, existing token-selection criteria based on local uncertainty, importance, or diversity do not directly determine whether changing the current selection im...

Jia Guo, Xiaohan Zhao, Changwang Liu et al. · 1 citation · ⚡1

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.