Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.
Identity matters in deep learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Gradient Descent Finds Global Minima for Generalizable Deep Neural Networks of Practical Sizes
Gradient descent on a network with a final hidden layer of width O(n) can interpolate any n-point dataset and reach a global optimum, and this linear rate is optimal.