Grams is an optimizer that forces the update direction to match the current gradient's sign while scaling by the magnitude of Adam's momentum, and the paper claims faster loss descent and global convergence.
signsgd: Compressed optimisation for non-convex problems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Grams: Gradient Descent with Adaptive Momentum Scaling
Grams is an optimizer that forces the update direction to match the current gradient's sign while scaling by the magnitude of Adam's momentum, and the paper claims faster loss descent and global convergence.