AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-$k$ number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval. ViSAR operates

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:32:15.665Z. This is not the publication date.