REVIEW 1 cited by
On the Ineffectiveness of Variance Reduced Optimization for Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The application of stochastic variance reduction to optimization has shown remarkable recent theoretical and practical success. The applicability of these techniques to the hard non-convex optimization problems encountered during training of modern deep neural networks is an open problem. We show that naive application of the SVRG technique and related approaches fail, and explore why.
Forward citations
Cited by 1 Pith paper
-
Towards Better Generalization: BP-SVRG in Training Deep Neural Networks
A sign-flipped SVRG variant called BP-SVRG adds stochastic-gradient noise instead of cancelling it, and empirically generalizes better than standard SVRG and often better than SGD on CIFAR and SVHN image classifiers.
Discussion (0). Continue with ORCID to comment.