Pith. sign in

Learning-Rate-Free Learning by D-Adaptation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

D-Adaptation is an approach to automatically setting the learning rate which asymptotically achieves the optimal rate of convergence for minimizing convex Lipschitz functions, with no back-tracking or line searches, and no additional function value or gradient evaluations per step. Our approach is the first hyper-parameter free method for this class without additional multiplicative log factors in the convergence rate. We present extensive experiments for SGD and Adam variants of our method, where the method automatically matches hand-tuned learning rates across more than a dozen diverse machine learning problems, including large-scale vision and language problems. An open-source implementation is available.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

LightSAM: Parameter-Agnostic Sharpness-Aware Minimization

cs.LG · 2025-05-30 · reject · novelty 6.0

An adaptive SAM variant using AdaGrad and Adam steps for both perturbation and update is claimed to converge at O(ln T / T^{1/4}) without tuning, but the Adam version still needs decaying hyperparameters and the proof contains an incorrect inequality.

citing papers explorer

Showing 1 of 1 citing paper.

  • LightSAM: Parameter-Agnostic Sharpness-Aware Minimization cs.LG · 2025-05-30 · reject · none · ref 14 · internal anchor

    An adaptive SAM variant using AdaGrad and Adam steps for both perturbation and update is claimed to converge at O(ln T / T^{1/4}) without tuning, but the Adam version still needs decaying hyperparameters and the proof contains an incorrect inequality.