Pith. sign in

REVIEW 1 cited by

Adaptive norms for deep learning with regularized Newton methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.09201 v4 pith:YRG3NEUN submitted 2019-05-22 cs.LG stat.ML

classification cs.LGstat.ML
keywords methodsadaptivenewtontermsbackpropagationsconstraintscounterpartellipsoidal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate the use of regularized Newton methods with adaptive norms for optimizing neural networks. This approach can be seen as a second-order counterpart of adaptive gradient methods, which we here show to be interpretable as first-order trust region methods with ellipsoidal constraints. In particular, we prove that the preconditioning matrix used in RMSProp and Adam satisfies the necessary conditions for provable convergence of second-order trust region methods with standard worst-case complexities on general non-convex objectives. Furthermore, we run experiments across different neural architectures and datasets to find that the ellipsoidal constraints constantly outperform their spherical counterpart both in terms of number of backpropagations and asymptotic loss value. Finally, we find comparable performance to state-of-the-art first-order methods in terms of backpropagations, but further advances in hardware are needed to render Newton methods competitive in terms of computational time.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stacey: Promoting Stochastic Steepest Descent via Accelerated $\ell_p$-Smooth Nonconvex Optimization

    cs.LG 2025-06 reject novelty 5.0 of 10

    STACEY is a new ℓ_p steepest descent optimizer with primal-dual interpolation; its convergence theory covers only the unaccelerated base algorithm, and its empirical gains rely on grid-searched hyperparameters.

Pith tools