REVIEW 1 cited by
Speed learning on the fly
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The practical performance of online stochastic gradient descent algorithms is highly dependent on the chosen step size, which must be tediously hand-tuned in many applications. The same is true for more advanced variants of stochastic gradients, such as SAGA, SVRG, or AdaGrad. Here we propose to adapt the step size by performing a gradient descent on the step size itself, viewing the whole performance of the learning trajectory as a function of step size. Importantly, this adaptation can be computed online at little cost, without having to iterate backward passes over the full data.
Forward citations
Cited by 1 Pith paper
-
A Learn-to-Optimize Approach for Coordinate-Wise Step Sizes for Quasi-Newton Methods
An LSTM-based model predicts coordinate-wise step sizes for BFGS and reports faster convergence while claiming theoretical guarantees that are only partially met.
Discussion (0). Continue with ORCID to comment.