REVIEW 5 cited by
On Empirical Comparisons of Optimizers for Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Selecting an optimizer is a central step in the contemporary deep learning pipeline. In this paper, we demonstrate the sensitivity of optimizer comparisons to the hyperparameter tuning protocol. Our findings suggest that the hyperparameter search space may be the single most important factor explaining the rankings obtained by recent empirical comparisons in the literature. In fact, we show that these results can be contradicted when hyperparameter search spaces are changed. As tuning effort grows without bound, more general optimizers should never underperform the ones they can approximate (i.e., Adam should never perform worse than momentum), but recent attempts to compare optimizers either assume these inclusion relationships are not practically relevant or restrict the hyperparameters in ways that break the inclusions. In our experiments, we find that inclusion relationships between optimizers matter in practice and always predict optimizer comparisons. In particular, we find that the popular adaptive gradient methods never underperform momentum or gradient descent. We also report practical tips around tuning often ignored hyperparameters of adaptive gradient methods and raise concerns about fairly benchmarking optimizers for neural network training.
Forward citations
Cited by 5 Pith papers
-
Polylogarithmic-Weight Dicke States in QAC$^0$ and Arbitrary Symmetric States in QAC$^0_f$
Polylog-weight n-qubit Dicke states are preparable in QAC⁰, and weight-k Dicke states are QAC⁰-preparable iff FANOUT_k is in QAC⁰ (for k ≤ n/2).
-
HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation
VLAC-Cut-guided multi-robot HITL post-training reaches 80–95% success and 1.7–4.2× throughput over the base VLA, outperforming HITL-only under the same human budget.
-
Path Integral Optimiser: Global Optimisation via Neural Schr\"odinger-F\"ollmer Diffusion
A neural Schrödinger-Föllmer diffusion, trained like the Path Integral Sampler, is repurposed as a global optimizer, with new conditional convergence bounds and competitive results on small tasks only.
-
Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI
Ambivision is a new dataset of AI-generated animal optical illusions, and the paper claims that adding gaze and eye annotations as visible image features improves classification while exposing limits of pixel-based ex...
-
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
Newer, larger deep learning models show no consistent gains over older architectures for speech emotion recognition across two naturalistic benchmarks, with results sensitive to model selection and hyperparameters.
Discussion (0). Sign in to comment.