A last-iterate convex optimization bound is shown to predict the shape and optimal learning-rate ratios of constant-plus-cooldown (wsd) and cosine schedules for LLM training, with validated learning-rate transfer rules.
For validation set metrics, we display a running average over five epoch in thick to smoothen the plot, and the original data in thin
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
A last-iterate convex optimization bound is shown to predict the shape and optimal learning-rate ratios of constant-plus-cooldown (wsd) and cosine schedules for LLM training, with validated learning-rate transfer rules.