AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, while later steps refine representations over largely the same routing support. Motivated by this structure, we introduce WISE (Working-set Inference

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T01:22:21.678Z. This is not the publication date.