For a univariate depth-2 linear network, gradient descent converges linearly to a global minimum even at large step sizes, and reaches a flatter minimum than gradient flow by shrinking the parameter imbalance.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
For a univariate depth-2 linear network, gradient descent converges linearly to a global minimum even at large step sizes, and reaches a flatter minimum than gradient flow by shrinking the parameter imbalance.