Pith. sign in

REVIEW

Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.08597 v5 pith:YQ3R4WVF submitted 2019-10-18 stat.ML cs.LGmath.OCstat.ME

classification stat.MLcs.LGmath.OCstat.ME
keywords learningmethodoptimizationratestochasticbetterdetectionlocal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes SplitSGD, a new dynamic learning rate schedule for stochastic optimization. This method decreases the learning rate for better adaptation to the local geometry of the objective function whenever a stationary phase is detected, that is, the iterates are likely to bounce at around a vicinity of a local minimum. The detection is performed by splitting the single thread into two and using the inner product of the gradients from the two threads as a measure of stationarity. Owing to this simple yet provably valid stationarity detection, SplitSGD is easy-to-implement and essentially does not incur additional computational cost than standard SGD. Through a series of extensive experiments, we show that this method is appropriate for both convex problems and training (non-convex) neural networks, with performance compared favorably to other stochastic optimization methods. Importantly, this method is observed to be very robust with a set of default parameters for a wide range of problems and, moreover, can yield better generalization performance than other adaptive gradient methods such as Adam.

Discussion (0). Sign in to comment.

Pith tools