Pith. sign in

REVIEW 2 cited by

Stochastic Gradient Descent for Two-layer Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07670 v1 pith:DTZK5XBC submitted 2024-07-10 stat.ML cs.LG

classification stat.MLcs.LG
keywords neuralnetworksconvergencetwo-layerkerneloverparameterizedalgorithmdependence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a comprehensive study on the convergence rates of the stochastic gradient descent (SGD) algorithm when applied to overparameterized two-layer neural networks. Our approach combines the Neural Tangent Kernel (NTK) approximation with convergence analysis in the Reproducing Kernel Hilbert Space (RKHS) generated by NTK, aiming to provide a deep understanding of the convergence behavior of SGD in overparameterized two-layer neural networks. Our research framework enables us to explore the intricate interplay between kernel methods and optimization processes, shedding light on the optimization dynamics and convergence properties of neural networks. In this study, we establish sharp convergence rates for the last iterate of the SGD algorithm in overparameterized two-layer neural networks. Additionally, we have made significant advancements in relaxing the constraints on the number of neurons, which have been reduced from exponential dependence to polynomial dependence on the sample size or number of iterations. This improvement allows for more flexibility in the design and scaling of neural networks, and will deepen our theoretical understanding of neural network models trained with SGD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

    stat.ML 2026-06 unverdicted novelty 7.0 of 10

    The paper derives the first minimax-optimal excess population risk rates for gradient descent and stochastic gradient descent on over-parameterized DNNs by linking their dynamics to kernel methods under polynomial wid...

  2. Spectral Algorithms in Misspecified Regression: Convergence under Covariate Shift

    stat.ML 2025-09 accept novelty 6.0 of 10

    Weighted spectral algorithms under covariate shift achieve minimax rates for bounded density ratios and near-optimal rates under truncation, including misspecified targets.

Pith tools