AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

arXiv · AI, language, vision and robotics · article · Sep 20, 2026 · UTC

Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T09:51:33.063Z. This is not the publication date.