{"id":"0066d8cc-8255-4116-b589-26926ed47f5b","arxiv_id":"2501.18965","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A last-iterate convex optimization bound is shown to predict the shape and optimal learning-rate ratios of constant-plus-cooldown (wsd) and cosine schedules for LLM training, with validated learning-rate transfer rules.","lead":"This paper shows that a theoretical bound from non-smooth convex optimization reproduces the shape of learning-rate curves seen when training large language models, including the sudden loss drop during cooldown. It uses the bound to derive practical rules for continued training and for transferring tuned learning rates across schedules.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The agreement is evaluated with G_t = 1, but Section 4.2 shows the cooldown drop disappears and the cosine/wsd LR ratio shifts for decaying G_t; the paper never measures G_t.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that verdict: the theoretical core (Theorem 3.1, Corollary 3.3, and the wsd bound in Theorem 3.4) is derived from an existing result and appears internally consistent, and the LR-transfer experiment on real 124M/210M runs plus the ImageNet SGD runs provide non-circular empirical checks. The most load-bearing soft spot is not the bare fact that the theory is for SGD on convex problems—the paper is candid about that in Section 6—but a specific, unverified modeling assumption inside the comparison: G_t is treated as constant. Section 4.2 is an explicit sensitivity analysis showing that the very phenomena used as evidence (sudden drop, cosine/wsd ratio) disappear when G_t decays like t^{-0.5} or t^{-1}. Since the empirical comparison is made against AdamW curves, the relevant gradient-norm profile is not even defined by the theory, making the claimed agreement fragile in a way that one measurement could resolve. Secondary issues—no seed variance, one disclosed post-hoc exclusion of cooldown fraction 0.6 in Fig. 21, and the WolframAlpha-backed identities in Lemma D.2—are real but do not change the conditional verdict. I would keep the CONDITIONAL verdict and add the gradient-norm measurement as the specific condition that would determine whether the central agreement is robust.","tokens_in":28522,"tokens_out":8301,"duration_ms":82633,"concrete_test":"Log per-step stochastic gradient norms (and, for comparison, AdamW's effective update norm) during a 124M or 210M Llama run over 50k steps, before and during cooldown. Fit G_t ≈ c·t^α over the pre-cooldown window (e.g., steps 1k–40k). If the fitted α is negative with magnitude ≥ 0.25, recompute Eq. (9) with that fitted G_t for both cosine and wsd and compare the bound shapes and the implied γ*_cos/γ*_wsd with Figs. 1–3 and the real sweep in Fig. 12b. A material mismatch would show the agreement is an artifact of the constant-G_t assumption; α ≈ 0 would instead support the paper's conjecture.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Figs. 1–3 evaluates Corollary 3.3 with G_t = D = 1, i.e., a constant bound on gradient norms. Section 4.2 shows this choice is not innocent: for G_t ∝ t^α with α = −0.5 or −1, the sudden wsd cooldown drop is no longer visible and the schedule comparison changes materially (Fig. 6). Thus the claimed agreement—the shape of the loss curves and the γ*_cos/γ*_wsd ≈ 2 ratio—is contingent on gradient norms not decaying over the run. The paper does not report gradient norms for any LLM, ImageNet, or OpenWebText2 run; it only conjectures in Section 7 that non-smoothness or gradient noise keeps them non-vanishing. Moreover, because the practical runs use AdamW, whose update is scale-invariant to the raw gradient, it is unclear what the SGD-bound's G_t should even be in the empirical comparison. This is a sharper version of the Section 6 limitation: even before worrying about non-convexity or SGD-vs-AdamW, the theory's qualitative predictions are sensitive to an unmeasured gradient-norm profile. If real G_t decays appreciably, the same bound that is claimed to explain the cooldown drop would predict no drop, and the transfer-rule match in Fig. 12a would lose its theoretical basis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the connection between a last-iterate suboptimality bound for non-smooth stochastic convex optimization (Defazio et al., 2023) and learning-rate schedules for large model training. It derives a bound for the wsd schedule, shows that the bound's shape over time resembles empirical validation-loss curves for cosine and wsd schedules, and uses the bound's minimizer to construct schedule-extension rules and to transfer optimal base learning-rates between schedules. The experiments include predictions for 124M and 210M Llama-type models, ImageNet ResNet50 with SGD, and OpenWebText2 language models. The paper also provides a mirror-descent extension in Appendix F and a PEP-based lower-bound analysis in Appendix B.4.","tokens_in":28742,"tokens_out":6658,"duration_ms":58656,"significance":"The paper offers a novel perspective on learning-rate scheduling by showing that predictions from a simple non-smooth convex bound can match empirical observations. The transfer rules and schedule-extension scheme are validated out-of-sample on real training runs, and the code is available, which is a genuine strength. However, the central link—from a worst-case SGD bound with an unmeasured gradient-norm profile to AdamW on non-convex transformers—is not fully established. Because the claimed 'surprising agreement' is the main thesis, this gap tempers the significance of the theoretical framing, even though the practical transfer rule (exp(0.7) factor) is a useful empirical contribution.","major_comments":[{"comment":"The conclusion in Sections 3.1 and 4.1 that the wsd cooldown drop and the cosine/wsd comparison are reproduced by the bound relies on the assumption G_t = 1 for all t. Figure 6 shows that the sudden drop disappears when G_t is proportional to t^alpha for alpha = -0.5 or -1, and the schedule comparison changes materially. Since no gradient-norm measurements are reported for any of the LLM, ImageNet, or OpenWebText2 runs, and since AdamW is invariant to the scale of the stochastic gradient, it remains unclear whether the empirical runs satisfy the required profile. This is a load-bearing assumption for the claimed agreement, and it should be tested by measuring G_t during training or replaced by a more robust theoretical condition.","section":"Section 4.2, Fig. 6"},{"comment":"The manuscript explicitly acknowledges that the theoretical results apply to SGD on convex objectives while all main experiments (Figs. 1, 10, 12 and the OpenWebText2/ImageNet replications) use AdamW or SGD on non-convex models. This gap is more than a caveat because the practical quantities used in Section 5, such as the schedule-reduction factor rho = 0.525 for T2 = 2T1 and the transfer factor exp(0.7) from Fig. 11, are derived from the convex bound. The evidence offered for transfer (the mirror-descent extension in Appendix F and references to SGD/Adam equivalence) is circumstantial. The paper should provide a direct empirical test of whether the bound's schedule-shape predictions hold for AdamW on a small transformer, or present the tuning rules as heuristics rather than theory-based rules.","section":"Section 6"},{"comment":"The agreement between the theoretical bound and empirical loss curves is assessed visually, without a quantitative measure. The bound has an arbitrary vertical scale (D = G = 1) and is a worst-case upper bound rather than a model of the loss trajectory. The authors should report a quantitative summary of the match, such as the correlation between the predicted and observed curves after optimally scaling the bound, or a defined feature (e.g., the drop height) with an error bar. Without such a measure, the 'surprisingly close match' claim is difficult to evaluate and could be confounded by the many degrees of freedom in the schedules and the chosen base learning-rates.","section":"Section 3.1, Figs. 1–3"}],"minor_comments":[{"comment":"The word 'converegnce' in the caption is a typo and should be corrected to 'convergence'.","section":"Fig. 6 caption"},{"comment":"The phrase 'the bound of the expected gradient norms G1:T' should be 'the bound on the expected gradient norms G1:T'.","section":"Section 4.2"},{"comment":"The symbol '≾' is used without a definition; please specify that it denotes an asymptotic inequality as T tends to infinity.","section":"Theorem 3.4"},{"comment":"The statement 'we verified that changing the values of G, D, or T1 do not affect the result' is not supported by a figure or table; a one-line explanation of the multiplicative scaling would be clearer.","section":"Section 5.1"},{"comment":"The description of the grid for base learning-rate gamma and cooldown fraction c does not mention the number of seeds or run-to-run variance, so the fitted optimum gamma*(c) in Fig. 12a has no uncertainty estimate.","section":"Appendix B.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is interesting and the practical transfer rule is a valuable empirical contribution. The main risk is that the 'surprising agreement' may depend on the specific unmeasured choice G_t = 1, as shown by the sensitivity analysis in Fig. 6. Given that the paper is already at a high standard of presentation and reproducibility, I think major_revision is appropriate: the authors should measure gradient norms, add a direct SGD-vs-AdamW comparison or reframe the tuning rules as heuristics, and include a quantitative measure of the shape agreement. These changes are feasible within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, honest paper that deserves a serious referee. The new content is real—the wsd bound in Theorem 3.4 (the log-free cooldown benefit) and the LR-transfer curves (the exp(0.7) factor from 20% cooldown to linear decay) are not in the cited literature, and the authors validate the transfer predictions on fresh 124M/210M runs rather than tuning constants to fit. The derivations in the appendix check out; the PEP lower bounds are a nice touch, and the ImageNet/OpenWebText2 experiments broaden the claim beyond one setup.\n\nThe soft spots are real but not fatal. First, the stress-test concern about G_t is legitimate and sharper than the paper's own Section 6. In Section 4.2 the authors show that if gradient norms decay like t^{-0.5} or t^{-1}, the characteristic wsd cooldown drop disappears from the bound, and the cosine/wsd LR ratio shifts. Yet they never measure gradient norms on any of the LLM, ImageNet, or OpenWebText2 runs. Since the practical optimizer is AdamW, whose updates are invariant to raw gradient scale, it is genuinely unclear what G_t should be in the empirical comparison. The paper conjectures in Section 7 that non-smoothness or gradient noise keeps gradient norms non-vanishing, but that is a conjecture, and it is load-bearing for the central shape-matching claim. This is addressable: measure per-layer gradient norms over a run, or test the schedule-shape predictions with SGD on a controlled non-convex problem.\n\nSecond, the empirical gains are small (about 0.01 validation loss) and reported without seed variance; the significance argument via scaling laws helps but is not a substitute for variance estimates. One run is excluded post hoc in Appendix B.6 with explanation, which is honest but should be reported in the main text.\n\nThat said, the central bound and transfer computations are correct as formal statements, and the out-of-sample validation is genuinely non-circular. The paper never overclaims—the limitations section names the SGD-to-AdamW and convex-to-non-convex bridges directly. I would send it to review. The main ask for revision: measure or bound G_t, add multi-seed variance for the headline experiments, and discuss the excluded run in the main text.","headline":"Solid, honest scheduling paper with real new results; the agreement claim needs one measured quantity (G_t) before it fully lands.","tokens_in":29420,"tokens_out":1945,"would_cite":true,"duration_ms":18475,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A last-iterate bound from non-smooth convex optimization, evaluated at its optimal base learning rate, reproduces the empirical loss curves of cosine and wsd schedules in LLM pretraining and yields transferable learning-rate rules.","keywords":["learning-rate schedules","warmup-stable-decay","cooldown","last-iterate convergence","non-smooth convex optimization","stochastic gradient descent","language model pretraining","learning-rate transfer"],"falsifier":"Train a sufficiently large model, such as a 1B-parameter transformer, with AdamW under cosine and wsd using bound-minimizing base rates, and measure the ratio $\\gamma^\\star_{\\text{cosine}}/\\gamma^\\star_{\\text{wsd}}$ and the size of the cooldown drop; if the ratio departs from roughly 2, the claimed agreement is scale- or optimizer-specific.","tokens_in":28243,"feed_emoji":"📉","tokens_out":8073,"duration_ms":68419,"temperature":0.7,"pith_summary":"This paper tries to establish that the shapes of learning-rate schedules used in large-model pretraining—cosine and warmup-stable-decay (wsd), including the sudden loss drop when cooldown begins—are reproduced by a worst-case bound from non-smooth stochastic convex optimization. The bound is for the last iterate of SGD, and once its base learning rate is tuned to minimize the bound, the theoretical curve tracks empirical validation-loss curves for transformer training. The paper derives a wsd-specific version of the bound and shows that cooldown removes logarithmic terms, giving a mechanism for the practical benefit of cooldown. It also converts the bound into tuning rules: the optimal base learning rate scales as the inverse square root of the horizon, cosine needs roughly twice the wsd learning rate, and an optimal rate for one cooldown fraction can be transferred to another. Correct, this would mean schedule design and learning-rate transfer for large models can be guided by a first-principles bound rather than by trial and error alone.","feed_headline":"A convex bound explains the cooldown loss drop","feed_subtitle":"One worst-case SGD bound reproduces cosine and wsd loss shapes and yields transferable learning rates.","key_machinery":"The load-bearing object is the last-iterate schedule bound from Theorem 3.1 (Eq. 6): for iterates $x_{t+1}=x_t-\\gamma\\eta_t g_t$ with convex losses, it upper-bounds $\\mathbb{E}[f(x_T)-f(x^\\star)]$ in terms of the initial distance $D$, gradient-norm bounds $G_t$, and the schedule $(\\eta_t)$, separated from a base rate $\\gamma$. Corollary 3.3 minimizes the bound over $\\gamma$, giving $\\gamma^\\star=\\sqrt{T_1/T_2}$ and the plug-in bound $2\\sqrt{T_1T_2}$. The paper then evaluates these for cosine and for wsd, which is constant until $T_0$ and then decays linearly; the wsd calculation is what makes cooldown visible as the disappearance of logarithmic terms, and the bound-minimizing rate is what supplies the learning-rate transfer rules.","core_discovery":"The central claim is that the last-iterate suboptimality bound for SGD on convex Lipschitz objectives, evaluated at the base learning rate that minimizes it, reproduces the empirical loss curves of cosine and wsd schedules in LLM pretraining. For wsd the paper derives a closed-form bound whose cooldown phase eliminates the logarithmic factor present for constant schedules, explaining the practical cooldown benefit. The same bound predicts that the optimal base learning rate decays like $T^{-1/2}$, that cosine's optimal rate is roughly twice wsd's, and that the ratio of optimal rates across cooldown fractions is stable across horizons; these predictions match re-analysis of real training runs and yield transfer factors such as $\\gamma^\\star(1)\\approx e^{0.7}\\gamma^\\star(0.2)$ for linear cooldown. The paper also shows the drop during cooldown appears in upper bounds, worst-case lower bounds computed by semidefinite programming, and a two-dimensional non-smooth convex problem, supporting the claim that the phenomenon is not architecture-specific.","pith_inferences":["Editorial inference: if the bound remains predictive beyond 210M parameters, schedule experiments could be cheaply screened against the theoretical curve before spending GPU hours, with the bound serving as a prior for candidate schedules.","Editorial inference: the paper's analysis suggests a testable decomposition—the cooldown drop is tied to non-vanishing gradient norms; monitoring $\\mathbb{E}\\|g_t\\|^2$ during pretraining could predict whether wsd's drop will be sharp from that curve alone.","Editorial inference: the mirror-descent extension in the appendix hints that sign-descent-like preconditioning may admit the same last-iterate bound, which would connect the SGD-based theory to Adam's success without invoking convexity of the full network.","Editorial inference: because the bound's schedule-shape predictions are invariant to the scale of $G$ and $D$, a direct test is to vary batch size or loss scaling and check whether the optimal rate ratios remain unchanged."],"forward_implications":["If the bound is the right testbed, a fully tuned base learning rate makes linear decay the optimal schedule among the studied classes, so the optimal cooldown fraction is one.","Continued training can be scheduled by theory: after extending a wsd run from $T_1$ to $T_2$, decreasing the schedule by a computed factor such as $\\rho=0.525$ for $T_2=2T_1$ keeps the bound close to a freshly tuned linear-decay run.","Learning-rate transfer across schedules becomes a calculable multiplier: with the optimal rate for 20% linear cooldown known, the linear-decay rate is $\\gamma^\\star(1)\\approx e^{0.7}\\gamma^\\star(0.2)$, avoiding a new sweep.","The same logic explains why cosine's cycle length of one is optimal, matching the empirical recommendation for language-model pretraining.","The improvement from adapted continued training is worth roughly 6-8% more tokens by the paper's scaling-law estimate, corresponding to about 0.01 validation loss for the 124M and 210M models."],"supporting_citations":[{"why":"Supplies the last-iterate schedule bound (Thm. 10) that is the paper's backbone, along with the linear-decay schedule comparison.","marker":"Defazio et al. (2023)"},{"why":"Provides the empirical wsd and cosine loss curves on Llama-type models that the theoretical bound is compared against, and the cooldown-drop observation.","marker":"Hägele et al. (2024)"},{"why":"Defines the cosine schedule that is one of the two central schedule classes.","marker":"Loshchilov & Hutter (2017)"},{"why":"Gives the standard cosine cycle-length recommendation and the scaling law used to value the 0.01 loss improvement.","marker":"Hoffmann et al. (2022)"},{"why":"Shows linear decay matches the worst-case lower bound for the last iterate, anchoring the optimality claims for cooldown fraction one.","marker":"Zamani & Glineur (2023)"},{"why":"Shows SGD on diagonal networks is equivalent to mirror descent, used to argue the SGD-based result can extend toward practical optimizers.","marker":"Even et al. (2023)"},{"why":"Supplies the semidefinite-programming tool used to compute lower bounds that confirm the cooldown drop in worst-case examples.","marker":"Goujaud et al. (2024)"},{"why":"Provides the empirical observation that cosine's optimal base rate is roughly twice wsd's, which the theoretical ratio reproduces.","marker":"Shen et al. (2024)"}],"fun_headline_variants":["Convex bound predicts cooldown loss drop","Transferable LLM rates from a convex bound","Why cooldown works: convex optimization theory","LLM cooldown loss drop explained by bound","Optimal learning rates from convex theory"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, flagged by the paper itself in Section 6, is that the schedule-shape behavior of AdamW on non-convex transformer training is governed by a worst-case bound proven for SGD on convex Lipschitz objectives; if that transfer fails at larger scales or for AdamW-specific dynamics, the central agreement and the derived tuning rules lose their foundation.","fun_headline_variants_meta":{"raw":{"variants":["Convex bound predicts cooldown loss drop","Transferable LLM rates from a convex bound","Why cooldown works: convex optimization theory","LLM cooldown loss drop explained by bound","Optimal learning rates from convex theory"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1652,"prompt_tokens":860,"completion_tokens":792,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":732}},"tokens_in":476,"tokens_out":792,"duration_ms":7348,"temperature":1.0,"reasoning_tokens":732,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T21:51:28.263831+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a sufficiently large model, such as a 1B-parameter transformer, with AdamW under cosine and wsd using bound-minimizing base rates, and measure the ratio $\\gamma^\\star_{\\text{cosine}}/\\gamma^\\star_{\\text{wsd}}$ and the size of the cooldown drop; if the ratio departs from roughly 2, the claimed agreement is scale- or optimizer-specific.","supporting_citations":[],"review_version":1}