REVIEW 4 cited by
Disentangling Adaptive Gradient Methods from Learning Rates
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate several confounding factors in the evaluation of optimization algorithms for deep learning. Primarily, we take a deeper look at how adaptive gradient methods interact with the learning rate schedule, a notoriously difficult-to-tune hyperparameter which has dramatic effects on the convergence and generalization of neural network training. We introduce a "grafting" experiment which decouples an update's magnitude from its direction, finding that many existing beliefs in the literature may have arisen from insufficient isolation of the implicit schedule of step sizes. Alongside this contribution, we present some empirical and theoretical retrospectives on the generalization of adaptive gradient methods, aimed at bringing more clarity to this space.
Forward citations
Cited by 4 Pith papers
-
Gradient Methods with Online Scaling Part I. Theoretical Foundations
Online scaled gradient methods adapt matrix step sizes via online learning, match the best fixed step size asymptotically, and achieve non-asymptotic superlinear convergence on smooth strongly convex problems.
-
MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning
Two-sided diagonal preconditioning before Newton-Schulz orthogonalization, plus an adaptive scalar stepsize, improves Muon's GPT-2 pretraining loss at nearly unchanged cost.
-
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
A KL-divergence view of Shampoo and SOAP yields KL-Shampoo and KL-SOAP, with KL-Shampoo outperforming both baselines in LLM pretraining without Adam grafting.
-
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
SGD with momentum can match Adam's performance in language modeling when trained with small batches and careful tuning, a result that contradicts several popular explanations for the optimizer gap.
Discussion (0). Continue with ORCID to comment.