Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.
Bounds on the number of hidden neur ons in multilayer perceptrons,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes
Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.