Pith. sign in

REVIEW 1 cited by

Step-size Adaptation Using Exponentiated Gradient Updates

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00145 v1 pith:76OP5KMH submitted 2022-01-31 cs.LG

classification cs.LG
keywords gradientstep-sizescaleapproachgainupdatesbeenexponentiated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Optimizers like Adam and AdaGrad have been very successful in training large-scale neural networks. Yet, the performance of these methods is heavily dependent on a carefully tuned learning rate schedule. We show that in many large-scale applications, augmenting a given optimizer with an adaptive tuning method of the step-size greatly improves the performance. More precisely, we maintain a global step-size scale for the update as well as a gain factor for each coordinate. We adjust the global scale based on the alignment of the average gradient and the current gradient vectors. A similar approach is used for updating the local gain factors. This type of step-size scale tuning has been done before with gradient descent updates. In this paper, we update the step-size scale and the gain variables with exponentiated gradient updates instead. Experimentally, we show that our approach can achieve compelling accuracy on standard models without using any specially tuned learning rate schedule. We also show the effectiveness of our approach for quickly adapting to distribution shifts in the data during training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Learn-to-Optimize Approach for Coordinate-Wise Step Sizes for Quasi-Newton Methods

    cs.LG 2024-11 conditional novelty 5.0 of 10

    An LSTM-based model predicts coordinate-wise step sizes for BFGS and reports faster convergence while claiming theoretical guarantees that are only partially met.

Pith tools