AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Attention-Aware Routing: Coupling Routing and Attention in MoEs

arXiv · AI, language, vision and robotics · article · Sep 17, 2026 · UTC

In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of the model's contextual state, disentangled from the hidden state. Keeping the base transformer entirely frozen, we train only the routing parameters, isolating routing as the sole variable. AAR improves GSM8K by +3.37 pp over a routing-only SFT basel

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:12:08.350Z. This is not the publication date.