Pith. sign in

REVIEW 1 cited by

On the different regimes of Stochastic Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.10688 v4 pith:WNZRLVQF submitted 2023-09-19 cs.LG cond-mat.dis-nnstat.ML

classification cs.LGcond-mat.dis-nnstat.ML
keywords sizedescentdifferentgradientregimesstochastictemperaturebatch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Modern deep networks are trained with stochastic gradient descent (SGD) whose key hyperparameters are the number of data considered at each step or batch size $B$, and the step size or learning rate $\eta$. For small $B$ and large $\eta$, SGD corresponds to a stochastic evolution of the parameters, whose noise amplitude is governed by the ''temperature'' $T\equiv \eta/B$. Yet this description is observed to break down for sufficiently large batches $B\geq B^*$, or simplifies to gradient descent (GD) when the temperature is sufficiently small. Understanding where these cross-overs take place remains a central challenge. Here, we resolve these questions for a teacher-student perceptron classification model and show empirically that our key predictions still apply to deep networks. Specifically, we obtain a phase diagram in the $B$-$\eta$ plane that separates three dynamical phases: (i) a noise-dominated SGD governed by temperature, (ii) a large-first-step-dominated SGD and (iii) GD. These different phases also correspond to different regimes of generalization error. Remarkably, our analysis reveals that the batch size $B^*$ separating regimes (i) and (ii) scale with the size $P$ of the training set, with an exponent that characterizes the hardness of the classification problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A solvable teacher-student perceptron model predicts that the optimal fraction of anomaly examples in training is generally away from 50%, with a sharp crossover as training noise increases.

Pith tools