AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

Transformer-based large language models often suffer from inter-layer parameter redundancy, where functional transformations are redundantly learned across network depths. We propose CS-MoE, a novel Transformer architecture featuring cross-layer expert sharing to address this inefficiency. Deviating from the widely used Mixture-of-Experts (MoE) architecture that terminates each Transformer block with layer-isolated experts, CS-MoE combines layer-independent experts with concurrent access to a centralized, globally shared expert pool. This \textit{Global Experts Sharing} mechanism enables elast

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T18:02:20.432Z. This is not the publication date.