Pith. sign in

REVIEW 5 cited by

On Empirical Comparisons of Optimizers for Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.05446 v3 pith:HE5DMRNN submitted 2019-10-11 cs.LG stat.ML

classification cs.LGstat.ML
keywords optimizerscomparisonsgradienthyperparameterneveroptimizertuningadaptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Selecting an optimizer is a central step in the contemporary deep learning pipeline. In this paper, we demonstrate the sensitivity of optimizer comparisons to the hyperparameter tuning protocol. Our findings suggest that the hyperparameter search space may be the single most important factor explaining the rankings obtained by recent empirical comparisons in the literature. In fact, we show that these results can be contradicted when hyperparameter search spaces are changed. As tuning effort grows without bound, more general optimizers should never underperform the ones they can approximate (i.e., Adam should never perform worse than momentum), but recent attempts to compare optimizers either assume these inclusion relationships are not practically relevant or restrict the hyperparameters in ways that break the inclusions. In our experiments, we find that inclusion relationships between optimizers matter in practice and always predict optimizer comparisons. In particular, we find that the popular adaptive gradient methods never underperform momentum or gradient descent. We also report practical tips around tuning often ignored hyperparameters of adaptive gradient methods and raise concerns about fairly benchmarking optimizers for neural network training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Polylogarithmic-Weight Dicke States in QAC$^0$ and Arbitrary Symmetric States in QAC$^0_f$

    quant-ph 2026-04 unverdicted novelty 8.0 of 10

    Polylog-weight n-qubit Dicke states are preparable in QAC⁰, and weight-k Dicke states are QAC⁰-preparable iff FANOUT_k is in QAC⁰ (for k ≤ n/2).

  2. HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    VLAC-Cut-guided multi-robot HITL post-training reaches 80–95% success and 1.7–4.2× throughput over the base VLA, outperforming HITL-only under the same human budget.

  3. Path Integral Optimiser: Global Optimisation via Neural Schr\"odinger-F\"ollmer Diffusion

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A neural Schrödinger-Föllmer diffusion, trained like the Path Integral Sampler, is repurposed as a global optimizer, with new conditional convergence bounds and competitive results on small tasks only.

  4. Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI

    cs.CV 2025-05 reject novelty 6.0 of 10

    Ambivision is a new dataset of AI-generated animal optical illusions, and the paper claims that adding gaze and eye annotations as visible image features improves classification while exposing limits of pixel-based ex...

  5. Charting 15 years of progress in deep learning for speech emotion recognition: A replication study

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Newer, larger deep learning models show no consistent gains over older architectures for speech emotion recognition across two naturalistic benchmarks, with results sensitive to model selection and hyperparameters.

Pith tools