AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

arXiv · AI, language, vision and robotics · article · Aug 27, 2026 · UTC

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads (VRHs), a small subset of attention heads (about 1.7-2.6%) that are causally responsible for grounding text descriptions to image regions. To find them, we recast existing head-scoring methods under

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:32:02.028Z. This is not the publication date.