Beyond muon: Mud (momentum decorrelation) for faster transformer training.arXiv preprint arXiv:2603.17970, 2026

Ben S Southworth, Stephen Thomas · 2026 · arXiv 2603.17970

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra

cs.LG · 2026-05-23 · conditional · novelty 5.0

Muon optimizer outperforms AdamW in ViT training on two image datasets, with gains that depend on data augmentation strength and are linked to wider singular-value spread in QKV gradients and prevention of late-training mode collapse in MLP blocks.

citing papers explorer

Showing 1 of 1 citing paper after filters.

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra cs.LG · 2026-05-23 · conditional · none · ref 32
Muon optimizer outperforms AdamW in ViT training on two image datasets, with gains that depend on data augmentation strength and are linked to wider singular-value spread in QKV gradients and prevention of late-training mode collapse in MLP blocks.

Beyond muon: Mud (momentum decorrelation) for faster transformer training.arXiv preprint arXiv:2603.17970, 2026

fields

years

verdicts

representative citing papers

citing papers explorer