Pith. sign in

REVIEW 2 cited by

Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02322 v1 pith:F66G75C2 submitted 2024-07-02 cs.LG math.PR

classification cs.LGmath.PR
keywords stochasticdescentdifferentialdistributiondynamicsequationsgradientmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the dynamics of a continuous-time model of the Stochastic Gradient Descent (SGD) for the least-square problem. Indeed, pursuing the work of Li et al. (2019), we analyze Stochastic Differential Equations (SDEs) that model SGD either in the case of the training loss (finite samples) or the population one (online setting). A key qualitative feature of the dynamics is the existence of a perfect interpolator of the data, irrespective of the sample size. In both scenarios, we provide precise, non-asymptotic rates of convergence to the (possibly degenerate) stationary distribution. Additionally, we describe this asymptotic distribution, offering estimates of its mean, deviations from it, and a proof of the emergence of heavy-tails related to the step-size magnitude. Numerical simulations supporting our findings are also presented.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    Presents a self-normalized subsampling procedure for asymptotically valid confidence regions from SGD iterates under both finite and infinite variance assumptions.

  2. Joint Learning in the Gaussian Single Index Model

    cs.LG 2025-05 conditional novelty 6.0 of 10

    In Gaussian single-index models, joint gradient flow over direction and link function converges to the true regression function from either sign of initial alignment, with rate governed by the information exponent.

Pith tools