Skip to content

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

Aug 2026 · 1 citation · 71 references
Computer Science

TL;DR

An embodied scene question-answering (QA) interface in which vision never enters the language model is studied, showing the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope.

Abstract

End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model (LLM) answers. On a visible-scope-matched, occlusion-audited campus-robot benchmark, under a prospectively frozen criterion, the serialized interface, using detectors fine-tuned in-domain within each fold, outperforms a zero-shot VLM whose language model has the same 7B scale (0.7892 vs 0.7462), with a larger margin at 3B (0.7673 vs 0.6913). Preregistered decoupling experiments show the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope. The advantage grows as the reader shrinks to 1.5B and reverses at 0.5B, and a ground-truth oracle locates the reader-capability floor. Under matched task supervision the interfaces converge: a VLM fine-tuned with low-rank adaptation (LoRA) overtakes the zero-shot system but only ties an equally supervised text reader (0.8441 vs 0.8396, no statistically resolved difference), and both routes remain perception-bound. Reported perception parameters are comparable to those of the VLM's vision tower, and total compute is not smaller.

View source

Similar papers

Preprint Sep 2026

Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the froze...

Can Polat, Mustafa Kurban, E. Serpedin et al. · 0 citations
Preprint Aug 2026

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model must use visual evidence to make a complete decision whose parts jointly satisfy global c...

Ning-Xin Pan, Han-Yu Li, Ye-Hui Tang · 0 citations
Preprint Aug 2026

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

This paper argues that what a VLA needs is not the ability to generate language, but the ability to consume grounded language, and introduces a framework that endows a VLA with language competence through in-context post-training and an agentic tool-use interface.

Jia-Rui Yang, Wen Huang, Jia-Le Zhang et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.