REVIEW 2 cited by
Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study the dynamics of a continuous-time model of the Stochastic Gradient Descent (SGD) for the least-square problem. Indeed, pursuing the work of Li et al. (2019), we analyze Stochastic Differential Equations (SDEs) that model SGD either in the case of the training loss (finite samples) or the population one (online setting). A key qualitative feature of the dynamics is the existence of a perfect interpolator of the data, irrespective of the sample size. In both scenarios, we provide precise, non-asymptotic rates of convergence to the (possibly degenerate) stationary distribution. Additionally, we describe this asymptotic distribution, offering estimates of its mean, deviations from it, and a proof of the emergence of heavy-tails related to the step-size magnitude. Numerical simulations supporting our findings are also presented.
Forward citations
Cited by 2 Pith papers
-
Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance
Presents a self-normalized subsampling procedure for asymptotically valid confidence regions from SGD iterates under both finite and infinite variance assumptions.
-
Joint Learning in the Gaussian Single Index Model
In Gaussian single-index models, joint gradient flow over direction and link function converges to the true regression function from either sign of initial alignment, with rate governed by the information exponent.
Discussion (0). Sign in to comment.