A synthesis of approximation, optimization, and generalization theory arguing that gradient descent's implicit norm control on weight directions explains why overparameterized deep networks generalize.
However, our theoretical understanding of deep learning, and thus the ability of developing principled improvements, has lagged behind
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization
A synthesis of approximation, optimization, and generalization theory arguing that gradient descent's implicit norm control on weight directions explains why overparameterized deep networks generalize.