Pith. sign in

REVIEW 4 cited by

Disentangling Adaptive Gradient Methods from Learning Rates

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.11803 v1 pith:NRMTXU6E submitted 2020-02-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords adaptivegradientlearningmethodsgeneralizationscheduleaimedalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate several confounding factors in the evaluation of optimization algorithms for deep learning. Primarily, we take a deeper look at how adaptive gradient methods interact with the learning rate schedule, a notoriously difficult-to-tune hyperparameter which has dramatic effects on the convergence and generalization of neural network training. We introduce a "grafting" experiment which decouples an update's magnitude from its direction, finding that many existing beliefs in the literature may have arisen from insufficient isolation of the implicit schedule of step sizes. Alongside this contribution, we present some empirical and theoretical retrospectives on the generalization of adaptive gradient methods, aimed at bringing more clarity to this space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gradient Methods with Online Scaling Part I. Theoretical Foundations

    math.OC 2025-05 conditional novelty 7.0 of 10

    Online scaled gradient methods adapt matrix step sizes via online learning, match the best fixed step size asymptotically, and achieve non-asymptotic superlinear convergence on smooth strongly convex problems.

  2. MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Two-sided diagonal preconditioning before Newton-Schulz orthogonalization, plus an adaptive scalar stepsize, improves Muon's GPT-2 pretraining loss at nearly unchanged cost.

  3. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization

    stat.ML 2025-09 conditional novelty 6.0 of 10

    A KL-divergence view of Shampoo and SOAP yields KL-Shampoo and KL-SOAP, with KL-Shampoo outperforming both baselines in LLM pretraining without Adam grafting.

  4. Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SGD with momentum can match Adam's performance in language modeling when trained with small batches and careful tuning, a result that contradicts several popular explanations for the optimizer gap.

Pith tools