AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert utilization toward uniformity. Post-training inherits a different situation: the base router already encodes non-uniform expert co-activation structure, which a re-imposed uniformity objective flattens away. We show that downstream performance depends instead on holding this inherited routing softly, a principle we term soft router anchoring, and instantiate it as Router Prior Bias (RPB), a training-time bias that pulls the router logits toward a prior read off the frozen base router while

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:22:01.598Z. This is not the publication date.