Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.
The lower bound of the capacity for a neural network with multiple hidden layers,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes
Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.