SOURCE-LINKED INTELLIGENCE
Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining
Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw final iterate or a checkpoint average). A schedule that promotes optimization progress may differ from one that minimizes variation in the raw final iterate. Separating these choices creates an opportunity to maintain progress late in training while reducing variation in the returned model. To this end, we propose \emph{Terminal Shrinkage Averagi
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-21T23:19:17.000Z
First collected: 2026-09-23T04:21:13.910Z. This is not the publication date.