AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Capability-Routed Visual Retrieval and Evidence Threading for Long-Context Document Question Answering

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

Annual reports, diligence packs, and infographic dashboards bury numbers in page images: axes, cell grids, and footnotes that OCR pipelines flatten and that page-level visual retrievers still treat as interchangeable in-context examples. We keep a frozen Qwen2.5-VL-7B-Instruct generator and a ColPali / VisRAG-Ret page index, and insert three modules. A capability-aware visual router (CAVR) tags each retrieved page as text, table, chart, layout, or mixed and mixes specialist experts before generation. Weak-to-strong page selection (WSPS) distils a frozen 7B answerability teacher into a 3B selec

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:32:20.942Z. This is not the publication date.