AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

ARM: Attention with Routed-Memory for Learnable Sparse Control

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to manage the growing cache. In this paper, we propose Attention with Routed Memory (ARM) a novel KV caching structure that introduces a fully differentiable, fixed-size memory system organized as a hierarchical router. Via a Gumbel-Softmax, ARM learns to select memory slots and perform

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T08:01:43.213Z. This is not the publication date.