AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

arXiv · AI, language, vision and robotics · article · Aug 29, 2026 · UTC

End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model (LLM) answers. On a visible-scope-matched, occlusion-audited campus-robot benchmark, under a prospective

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T21:41:48.575Z. This is not the publication date.