Gradient descent iterates of finite-width multi-layer networks on single-index data obey a state evolution law, giving exact training/test error formulas and a data-driven test error estimator.
The LASSO risk for Gaussian matrices
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Precise gradient descent training dynamics for finite-width multi-layer neural networks
Gradient descent iterates of finite-width multi-layer networks on single-index data obey a state evolution law, giving exact training/test error formulas and a data-driven test error estimator.