AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three l

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T04:51:57.792Z. This is not the publication date.