Pith. sign in

REVIEW 2 cited by

Grad-GradaGrad? A Non-Monotone Adaptive Stochastic Gradient Method

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.06900 v1 pith:PAKMTQSS submitted 2022-06-14 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords methodratelearningadagraddecreasedenominatornon-monotoneaccumulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The classical AdaGrad method adapts the learning rate by dividing by the square root of a sum of squared gradients. Because this sum on the denominator is increasing, the method can only decrease step sizes over time, and requires a learning rate scaling hyper-parameter to be carefully tuned. To overcome this restriction, we introduce GradaGrad, a method in the same family that naturally grows or shrinks the learning rate based on a different accumulation in the denominator, one that can both increase and decrease. We show that it obeys a similar convergence rate as AdaGrad and demonstrate its non-monotone adaptation capability with experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization

    math.OC 2025-07 conditional novelty 7.0 of 10

    Accelerated GRAAL is the first adaptive first-order method that proves near-optimal accelerated complexity for convex L-smooth and (L0,L1)-smooth functions with geometric stepsize growth.

  2. Low-Rank Dependence Decomposition via Accelerated Symmetric Non-negative Matrix Factorization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Trace-reformulated SymNMF scales to n=10^6 on GPUs; five AdaGrad-family methods converge, with Block-SVRG AdaptGrow winning on flat TPDM spectra and full-batch AdaGrad on low-rank correlation spectra.

Pith tools