Pith. sign in

REVIEW 1 cited by

Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.02588 v2 pith:T7VRN7XZ submitted 2021-06-04 cs.LG math.APstat.ML

classification cs.LGmath.APstat.ML
keywords noisecontinuoustimealgorithmdescentflatgradientlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The representation of functions by artificial neural networks depends on a large number of parameters in a non-linear fashion. Suitable parameters of these are found by minimizing a 'loss functional', typically by stochastic gradient descent (SGD) or an advanced SGD-based algorithm. In a continuous time model for SGD with noise that follows the 'machine learning scaling', we show that in a certain noise regime, the optimization algorithm prefers 'flat' minima of the objective function in a sense which is different from the flat minimum selection of continuous time SGD with homogeneous noise.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Scale vectors in Pre-Norm LLMs aid optimization via preconditioning on linear layers rather than expressivity, and three lightweight modifications to them reduce terminal loss across model scales.

Pith tools