Pith. sign in

REVIEW 1 cited by

On the Generalization Benefit of Noise in Stochastic Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.15081 v1 pith:YPDCWFK5 submitted 2020-06-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords largestochasticbatchdescentgradientbatchesgeneralizationhyperparameter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have questioned this claim, arguing that this effect is simply a consequence of suboptimal hyperparameter tuning or insufficient compute budgets when the batch size is large. In this paper, we perform carefully designed experiments and rigorous hyperparameter sweeps on a range of popular models, which verify that small or moderately large batch sizes can substantially outperform very large batches on the test set. This occurs even when both models are trained for the same number of iterations and large batches achieve smaller training losses. Our results confirm that the noise in stochastic gradients can enhance generalization. We study how the optimal learning rate schedule changes as the epoch budget grows, and we provide a theoretical account of our observations based on the stochastic differential equation perspective of SGD dynamics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training

    cs.LG 2025-05 reject novelty 5.0 of 10

    SGD is said to minimize free energy, but the temperature is constructed from the data, making the validation circular.

Pith tools