A line-search-free adaptive rule for decentralized convex optimization, built on a new Lyapunov function and local secant curvature estimates.
Adaptive gradient descent without descent.arXiv preprint arXiv:1910.09529, 2019
8 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Develops constant-stepsize and auto-conditioned projected gradient methods plus stochastic variants that achieve new iteration complexity bounds for finding approximate stationary points in nonconvex smooth optimization.
SGD on multiclass cross-entropy loss alternates between curvature-driven oscillations and stable regimes but self-stabilizes to enable best-iterate convergence with large learning rates for linear and two-layer models.
Presents a model-based proximal framework for adaptive momentum in first-order optimizers by using a two-plane approximation of the objective to dynamically set the memory coefficient online.
AdaNAGED combines zeroth-order gradient-free training, automatic parameter adaptation, and LMO-based non-Euclidean geometry with claimed convergence guarantees, demonstrated on OPT-1.3B fine-tuning.
A Bregman proximal point method for CVaR minimization that uses dual-stage updates to generate importance sampling from the risk tail, with a claimed convergence proof for convex problems.
AAMD combines preconditioning, acceleration, and adaptivity in mirror descent using a Lyapunov budget to achieve O(1/k^2) rates under dual relative smoothness and bounded sublevel sets.
A proximal stochastic gradient method with variance reduction and adaptive steps is shown to converge strongly at rate O(sqrt(1/k)) for convex composite problems when the smooth term is Lipschitz continuous.
citing papers explorer
-
A Line-search-free Method for Adaptive Decentralized Optimization
A line-search-free adaptive rule for decentralized convex optimization, built on a new Lyapunov function and local secant curvature estimates.
-
Projected gradient methods for nonconvex and stochastic smooth optimization: new complexities and auto-conditioned stepsizes
Develops constant-stepsize and auto-conditioned projected gradient methods plus stochastic variants that achieve new iteration complexity bounds for finding approximate stationary points in nonconvex smooth optimization.
-
SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates
SGD on multiclass cross-entropy loss alternates between curvature-driven oscillations and stable regimes but self-stabilizes to enable best-iterate convergence with large learning rates for linear and two-layer models.
-
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
Presents a model-based proximal framework for adaptive momentum in first-order optimizers by using a two-plane approximation of the objective to dynamically set the memory coefficient online.
-
Zero-order Parameter-free Optimization for LMO-based Methods: Novel Approach for Efficient Fine-tuning
AdaNAGED combines zeroth-order gradient-free training, automatic parameter adaptation, and LMO-based non-Euclidean geometry with claimed convergence guarantees, demonstrated on OPT-1.3B fine-tuning.
-
Exponential Adaptive Smoothing and Importance Sampling for Optimization of the Conditional Value-at-Risk
A Bregman proximal point method for CVaR minimization that uses dual-stage updates to generate importance sampling from the risk tail, with a claimed convergence proof for convex problems.
-
Adaptive Accelerated Mirror Descent in Primal and Dual Spaces
AAMD combines preconditioning, acceleration, and adaptivity in mirror descent using a Lyapunov budget to achieve O(1/k^2) rates under dual relative smoothness and bounded sublevel sets.
-
A Proximal Stochastic Gradient Method with Adaptive Step Size and Variance Reduction for Convex Composite Optimization
A proximal stochastic gradient method with variance reduction and adaptive steps is shown to converge strongly at rate O(sqrt(1/k)) for convex composite problems when the smooth term is Lipschitz continuous.