Pith. sign in

The Spectral Dynamics and Noise Geometry of Muon

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Muon replaces a matrix gradient $G=U\Sigma V^\top$ by its polar factor $UV^\top$. This keeps the singular directions selected by the gradient, but makes the update spectrum flat. We study the optimization bias created by this operation. Under explicit alignment assumptions, we prove that the polar update is the one-step entropy-maximizing choice among bounded updates that use the gradient singular directions and do not adapt to the current weight spectrum. In an underdetermined regression model, we derive exact singular-value dynamics for continuous-time Muon and identify a measurement-dependent condition under which the normalized spectrum moves toward equal nonzero singular values. This geometry also rules out a common low-rank interpretation: at fixed Frobenius norm, Muon's distinguished state has a flat spectrum, whereas nuclear-norm minimization favors spectral concentration. Controlled matrix-sensing experiments separate the effect from simple gradient rescaling, show that norm-matched gradient descent does not reproduce Muon, and recover the predicted flattening trend across broad ablations. In small NanoGPT pretraining, Muon preserves stable rank, has a broad learning-rate plateau, and improves validation loss relative to AdamW; in a matched small-ViT control, the ranking reverses. The resulting picture is regime-dependent: Muon is not universally superior, but its flat-spectrum bias can help when many spectral directions need to remain active.

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

The Loss Does Not See the Basis, but Adam Does

cs.LG · 2026-08-05 · conditional · novelty 7.0

Whether an optimizer keeps gradient descent's low-rank bias in factored models is determined by its equivariance under orthogonal gauge rotations, and Adam and other coordinate-wise rules fail this test.

citing papers explorer

Showing 1 of 1 citing paper.

  • The Loss Does Not See the Basis, but Adam Does cs.LG · 2026-08-05 · conditional · none · ref 5 · internal anchor

    Whether an optimizer keeps gradient descent's low-rank bias in factored models is determined by its equivariance under orthogonal gauge rotations, and Adam and other coordinate-wise rules fail this test.