SOURCE-LINKED INTELLIGENCE
Accelerating Sharded Data Parallelism at Scale with Federated Learning
The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence. Foundation models (FMs) are a crucial example, requiring months-long training on thousands of cutting-edge GPUs. Sharded data parallelism (DP) is the dominant strategy to accelerate such computations by splitting data and models across multiple GPUs. However, it incurs prohibitive communication overhead when deployed at scale, particularly on multi-tier interconnects with heterogeneous performance. Inspired by the efficient communicatio
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-17T13:21:52.000Z
- arXiv · AI, language, vision and robotics · 2026-09-17T13:21:52.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.