Pith. sign in

Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 3 2025 1

roles

background 1

polarities

background 1

representative citing papers

Demystifying Manifold Constraints in LLM Pre-training

cs.LG · 2026-05-06 · unverdicted · novelty 6.0

Manifold constraints via the new MACRO optimizer independently bound activation scales and enforce rotational equilibrium in LLM pre-training, subsuming RMS normalization and decoupled weight decay while delivering competitive performance with convergence guarantees.

Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer

cs.LG · 2026-05-05 · unverdicted · novelty 4.0

Nora is a matrix optimizer that stabilizes weight norms and angular velocities through row-wise momentum projection onto the orthogonal complement of the weights while approximating structured preconditioning with O(mn) complexity and proven scalability.

citing papers explorer

Showing 4 of 4 citing papers.

  • Demystifying Manifold Constraints in LLM Pre-training cs.LG · 2026-05-06 · unverdicted · none · ref 59

    Manifold constraints via the new MACRO optimizer independently bound activation scales and enforce rotational equilibrium in LLM pre-training, subsuming RMS normalization and decoupled weight decay while delivering competitive performance with convergence guarantees.

  • mRadNet: A Compact Radar Object Detector with MetaFormer eess.SP · 2025-09-11 · unverdicted · none · ref 24

    mRadNet improves state-of-the-art radar object detection on the CRUW dataset while using the fewest parameters and lowest FLOPs among compared models.

  • Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cs.LG · 2026-06-24 · conditional · none · ref 28

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  • Nora: Normalized Orthogonal Row Alignment for Scalable Matrix Optimizer cs.LG · 2026-05-05 · unverdicted · none · ref 16

    Nora is a matrix optimizer that stabilizes weight norms and angular velocities through row-wise momentum projection onto the orthogonal complement of the weights while approximating structured preconditioning with O(mn) complexity and proven scalability.