Early-stopped gradient descent dominates ridge regression for all well-specified linear regression problems, and it dominates stochastic gradient descent whenever the covariance spectrum decays fast and continuously.
For valid generalization the size of the weights is more important than the size of the network
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
Early-stopped gradient descent dominates ridge regression for all well-specified linear regression problems, and it dominates stochastic gradient descent whenever the covariance spectrum decays fast and continuously.