Smooth deep networks trained by GD/SGD attain the minimax-optimal excess risk O(n^{-2β/(2β+γ)}) with polynomially large width, matching kernel methods under NTK source/capacity conditions.
On the rate of convergence of an over-parametrized deep neural network regression estimate learned by gradient descent.arXiv preprint arXiv:2504.03405, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent
Smooth deep networks trained by GD/SGD attain the minimax-optimal excess risk O(n^{-2β/(2β+γ)}) with polynomially large width, matching kernel methods under NTK source/capacity conditions.