A last-iterate convex optimization bound is shown to predict the shape and optimal learning-rate ratios of constant-plus-cooldown (wsd) and cosine schedules for LLM training, with validated learning-rate transfer rules.
2, but with Ωt from (11) The bound on the best-so-far bound has a very different shape of the last-iterate bound
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
A last-iterate convex optimization bound is shown to predict the shape and optimal learning-rate ratios of constant-plus-cooldown (wsd) and cosine schedules for LLM training, with validated learning-rate transfer rules.