Pith. sign in

REVIEW 3 cited by

Towards Quantifying the Preconditioning Effect of Adam

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07114 v1 pith:62IZ776N submitted 2024-02-11 cs.LG cs.NAmath.NAmath.OCstat.ML

classification cs.LGcs.NAmath.NAmath.OCstat.ML
keywords adamkappahessianconditionmathcaleffectnumberpreconditioning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

There is a notable dearth of results characterizing the preconditioning effect of Adam and showing how it may alleviate the curse of ill-conditioning -- an issue plaguing gradient descent (GD). In this work, we perform a detailed analysis of Adam's preconditioning effect for quadratic functions and quantify to what extent Adam can mitigate the dependence on the condition number of the Hessian. Our key finding is that Adam can suffer less from the condition number but at the expense of suffering a dimension-dependent quantity. Specifically, for a $d$-dimensional quadratic with a diagonal Hessian having condition number $\kappa$, we show that the effective condition number-like quantity controlling the iteration complexity of Adam without momentum is $\mathcal{O}(\min(d, \kappa))$. For a diagonally dominant Hessian, we obtain a bound of $\mathcal{O}(\min(d \sqrt{d \kappa}, \kappa))$ for the corresponding quantity. Thus, when $d < \mathcal{O}(\kappa^p)$ where $p = 1$ for a diagonal Hessian and $p = 1/3$ for a diagonally dominant Hessian, Adam can outperform GD (which has an $\mathcal{O}(\kappa)$ dependence). On the negative side, our results suggest that Adam can be worse than GD for a sufficiently non-diagonal Hessian even if $d \ll \mathcal{O}(\kappa^{1/3})$; we corroborate this with empirical evidence. Finally, we extend our analysis to functions satisfying per-coordinate Lipschitz smoothness and a modified version of the Polyak-\L ojasiewicz condition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety

    cs.LG 2026-07 conditional novelty 7.0 of 10

    HarmAlign localizes spectral curvature inflation to an estimated harmful activation subspace, blocking harmful fine-tuning within a first-order, fixed-architecture threat model while preserving benign trainability.

  2. Linear-Time User-Level DP-SCO via Robust Statistics

    cs.LG 2025-02 conditional novelty 7.0 of 10

    A linear-time algorithm using robust statistics achieves near-optimal excess risk for user-level private convex optimization under ℓ1/ℓ∞ geometry, up to an extra factor of ε.

  3. LightSAM: Parameter-Agnostic Sharpness-Aware Minimization

    cs.LG 2025-05 reject novelty 6.0 of 10

    An adaptive SAM variant using AdaGrad and Adam steps for both perturbation and update is claimed to converge at O(ln T / T^{1/4}) without tuning, but the Adam version still needs decaying hyperparameters and the proof...

Pith tools