Pith. sign in

REVIEW 2 cited by

Adaptive Gradient Descent without Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.09529 v2 pith:LSMPQUTB submitted 2019-10-21 math.OC cs.LGcs.NAmath.NAstat.ML

classification math.OCcs.LGcs.NAmath.NAstat.ML
keywords convexdescentadaptivefunctiongradientlocalmethodrules
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a strikingly simple proof that two rules are sufficient to automate gradient descent: 1) don't increase the stepsize too fast and 2) don't overstep the local curvature. No need for functional values, no line search, no information about the function except for the gradients. By following these rules, you get a method adaptive to the local geometry, with convergence guarantees depending only on the smoothness in a neighborhood of a solution. Given that the problem is convex, our method converges even if the global smoothness constant is infinity. As an illustration, it can minimize arbitrary continuously twice-differentiable convex function. We examine its performance on a range of convex and nonconvex problems, including logistic regression and matrix factorization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Line-search-free Method for Adaptive Decentralized Optimization

    math.OC 2026-05 unverdicted novelty 7.0 of 10

    New adaptive decentralized algorithms select stepsizes from local curvature estimates derived from a Lyapunov function, delivering sublinear convergence for convex problems and linear rates for strongly convex ones.

  2. AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent

    cs.LG 2025-05 conditional novelty 6.0 of 10

    AutoSGD runs three parallel SGD streams at nearby learning rates, uses paired noisy objective estimates to pick the winner, and is claimed to converge with little user tuning.

Pith tools