REVIEW 8 cited by
On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Conventional wisdom in deep learning states that increasing depth improves expressiveness but complicates optimization. This paper suggests that, sometimes, increasing depth can speed up optimization. The effect of depth on optimization is decoupled from expressiveness by focusing on settings where additional layers amount to overparameterization - linear neural networks, a well-studied model. Theoretical analysis, as well as experiments, show that here depth acts as a preconditioner which may accelerate convergence. Even on simple convex problems such as linear regression with $\ell_p$ loss, $p>2$, gradient descent can benefit from transitioning to a non-convex overparameterized objective, more than it would from some common acceleration schemes. We also prove that it is mathematically impossible to obtain the acceleration effect of overparametrization via gradients of any regularizer.
Forward citations
Cited by 8 Pith papers
-
A Theory of Saddle Escape in Deep Nonlinear Networks
Derives exact norm-imbalance identity for deep nonlinear nets, classifying activations into four classes and yielding escape time law τ★ = Θ(ε^{-(r-2)}) governed by bottleneck depth r.
-
Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks
Dead-Direction Conditioners provide gauge-equivariant preconditioning by conditioning optimizer state on symmetry orbits, yielding improved resistance to over-training collapse and higher detection of dead directions ...
-
Faster Query-Key Learning Sharpens Attention in Self-Attention Models
Faster query-key learning relative to output-value learning sharpens attention onto task-relevant tokens at comparable prediction performance, derived from gradient-flow dynamics and shown on synthetic and real tasks.
-
What Does Flow Matching Bring To TD Learning?
Flow matching critics outperform monolithic ones in RL by 2x performance and 5x sample efficiency via test-time error recovery through integration and multi-point velocity supervision that preserves feature plasticity.
-
The learnability scaling of quantum states: restricted Boltzmann machines
To reproduce the ground-state energy of a one-dimensional transverse-field Ising chain near its critical point, a restricted Boltzmann machine needs a number of weights that grows as the square of the number of qubits...
-
Towards Better Generalization: BP-SVRG in Training Deep Neural Networks
A sign-flipped SVRG variant called BP-SVRG adds stochastic-gradient noise instead of cancelling it, and empirically generalizes better than standard SVRG and often better than SGD on CIFAR and SVHN image classifiers.
-
Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
A single multilingual NMT model for 103 languages trained on 25B examples demonstrates transfer learning benefits for low-resource languages.
-
Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
Provides Hessian-based theoretical characterizations of SGD dynamics and a scale-invariant generalization bound for deep nets, backed by experiments on synthetic data, MNIST, and CIFAR-10.
Discussion (0). Continue with ORCID to comment.