REVIEW 2 cited by
Grad-GradaGrad? A Non-Monotone Adaptive Stochastic Gradient Method
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The classical AdaGrad method adapts the learning rate by dividing by the square root of a sum of squared gradients. Because this sum on the denominator is increasing, the method can only decrease step sizes over time, and requires a learning rate scaling hyper-parameter to be carefully tuned. To overcome this restriction, we introduce GradaGrad, a method in the same family that naturally grows or shrinks the learning rate based on a different accumulation in the denominator, one that can both increase and decrease. We show that it obeys a similar convergence rate as AdaGrad and demonstrate its non-monotone adaptation capability with experiments.
Forward citations
Cited by 2 Pith papers
-
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
Accelerated GRAAL is the first adaptive first-order method that proves near-optimal accelerated complexity for convex L-smooth and (L0,L1)-smooth functions with geometric stepsize growth.
-
Low-Rank Dependence Decomposition via Accelerated Symmetric Non-negative Matrix Factorization
Trace-reformulated SymNMF scales to n=10^6 on GPUs; five AdaGrad-family methods converge, with Block-SVRG AdaptGrow winning on flat TPDM spectra and full-batch AdaGrad on low-rank correlation spectra.
Discussion (0). Continue with ORCID to comment.