{"slug": "terminal-shrinkage-averaging-reveals-a-schedule-estimator-interaction-in-llm", "title": "Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining", "summary": "A new arXiv paper (2609.25482v1) proposes Terminal Shrinkage Averaging (TSA), a pretraining estimator that interpolates between the raw final iterate and the average of recent checkpoints to balance recent optimization progress against terminal variation. The authors analyze how TSA shifts the preferred terminal learning-rate schedule under a local quadratic approximation and test the interaction in controlled NanoChat experiments, then show the combined schedule and estimator improve validation quality on depth-22 NanoChat. A qualifying time-to-GPT-2 run also finished faster than the public baseline used in the experiments, which the authors call preliminary evidence of benchmark acceleration.", "body_md": "arXiv:2609.25482v1 Announce Type: new \nAbstract: Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw final iterate or a checkpoint average). A schedule that promotes optimization progress may differ from one that minimizes variation in the raw final iterate. Separating these choices creates an opportunity to maintain progress late in training while reducing variation in the returned model. To this end, we propose \\emph{Terminal Shrinkage Averaging (TSA)}, which interpolates between the raw final iterate and the average of recent checkpoints to balance recent progress against terminal variation. We analyze how TSA changes the preferred terminal learning-rate schedule under a local quadratic approximation and test this interaction through a sequence of controlled NanoChat experiments. Finally, we demonstrate that the resulting gains transfer to depth-22 NanoChat, where the combined schedule and estimator improve validation quality. A qualifying time-to-GPT-2 run also finishes faster than the public baseline used in our experiments, providing preliminary evidence of benchmark acceleration.", "url": "https://wpnews.pro/news/terminal-shrinkage-averaging-reveals-a-schedule-estimator-interaction-in-llm", "canonical_source": "https://www.machinebrief.com/news/terminal-shrinkage-averaging-reveals-a-schedule-estimator-in-v244", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:54:48.461007+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["Terminal Shrinkage Averaging", "NanoChat", "GPT-2", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/terminal-shrinkage-averaging-reveals-a-schedule-estimator-interaction-in-llm", "markdown": "https://wpnews.pro/news/terminal-shrinkage-averaging-reveals-a-schedule-estimator-interaction-in-llm.md", "text": "https://wpnews.pro/news/terminal-shrinkage-averaging-reveals-a-schedule-estimator-interaction-in-llm.txt", "jsonld": "https://wpnews.pro/news/terminal-shrinkage-averaging-reveals-a-schedule-estimator-interaction-in-llm.jsonld"}}