Pith. sign in

The Crucial Role of Normalization in Sharpness-Aware Minimization

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Sharpness-Aware Minimization (SAM) is a recently proposed gradient-based optimizer (Foret et al., ICLR 2021) that greatly improves the prediction performance of deep neural networks. Consequently, there has been a surge of interest in explaining its empirical success. We focus, in particular, on understanding the role played by normalization, a key component of the SAM updates. We theoretically and empirically study the effect of normalization in SAM for both convex and non-convex functions, revealing two key roles played by normalization: i) it helps in stabilizing the algorithm; and ii) it enables the algorithm to drift along a continuum (manifold) of minima -- a property identified by recent theoretical works that is the key to better performance. We further argue that these two properties of normalization make SAM robust against the choice of hyper-parameters, supporting the practicality of SAM. Our conclusions are backed by various experiments.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

unclear 1

representative citing papers

LightSAM: Parameter-Agnostic Sharpness-Aware Minimization

cs.LG · 2025-05-30 · reject · novelty 6.0

An adaptive SAM variant using AdaGrad and Adam steps for both perturbation and update is claimed to converge at O(ln T / T^{1/4}) without tuning, but the Adam version still needs decaying hyperparameters and the proof contains an incorrect inequality.

citing papers explorer

Showing 1 of 1 citing paper.

  • LightSAM: Parameter-Agnostic Sharpness-Aware Minimization cs.LG · 2025-05-30 · reject · none · ref 8 · internal anchor

    An adaptive SAM variant using AdaGrad and Adam steps for both perturbation and update is claimed to converge at O(ln T / T^{1/4}) without tuning, but the Adam version still needs decaying hyperparameters and the proof contains an incorrect inequality.