Pith. sign in

REVIEW 1 cited by

On the connections between optimization algorithms, Lyapunov functions, and differential equations: theory and insights

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.08658 v3 pith:NGUGNZDV submitted 2023-05-15 math.OC cs.NAmath.NAstat.ML

On the connections between optimization algorithms, Lyapunov functions, and differential equations: theory and insights

classification math.OC cs.NAmath.NAstat.ML
keywords algorithmsconvergencedifferentialequationfunctionsoptimizationpolyakable
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We revisit the general framework introduced by Fazylab et al. (SIAM J. Optim. 28, 2018) to construct Lyapunov functions for optimization algorithms in discrete and continuous time. For smooth, strongly convex objective functions, we relax the requirements necessary for such a construction. As a result we are able to prove for Polyak's ordinary differential equations and for a two-parameter family of Nesterov algorithms rates of convergence that improve on those available in the literature. We analyse the interpretation of Nesterov algorithms as discretizations of the Polyak equation. We show that the algorithms are instances of Additive Runge-Kutta integrators and discuss the reasons why most discretizations of the differential equation do not result in optimization algorithms with acceleration. We also introduce a modification of Polyak's equation and study its convergence properties. Finally we extend the general framework to the stochastic scenario and consider an application to random algorithms with acceleration for overparameterized models; again we are able to prove convergence rates that improve on those in the literature.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Adaptive Momentum and Nonlinear Damping for Neural Network Training

    cs.LG 2026-01 conditional novelty 6.0

    Per-parameter kinetic-energy friction and cubic damping in momentum optimizers close much of the Adam–mSGD gap on transformer training, with deterministic exponential-convergence guarantees for strongly convex losses.