Gradient descent in a 1D shallow network acts like a shrinkage operator on the Jacobian's singular values, so the learning rate and number of iterations explicitly set the spectral bandwidth for monotonic activations.
Stability and generalization of learning algorithms that converge to global 8 optima
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gradient Descent as a Shrinkage Operator for Spectral Bias
Gradient descent in a 1D shallow network acts like a shrinkage operator on the Jacobian's singular values, so the learning rate and number of iterations explicitly set the spectral bandwidth for monotonic activations.