Muon-accelerated attention distillation for real-time edge synthesis via optimized latent diffusion.arXiv preprint arXiv:2504.08451, 2025

Weiye Chen, Qingen Zhu, Qian Long · 2025 · arXiv 2504.08451

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra

cs.LG · 2026-05-23 · conditional · novelty 5.0

Muon optimizer outperforms AdamW in ViT training on two image datasets, with gains that depend on data augmentation strength and are linked to wider singular-value spread in QKV gradients and prevention of late-training mode collapse in MLP blocks.

citing papers explorer

Showing 1 of 1 citing paper.

Muon in Vision Transformers: Optimizer-Recipe Interactions and Gradient Spectra cs.LG · 2026-05-23 · conditional · none · ref 8
Muon optimizer outperforms AdamW in ViT training on two image datasets, with gains that depend on data augmentation strength and are linked to wider singular-value spread in QKV gradients and prevention of late-training mode collapse in MLP blocks.

Muon-accelerated attention distillation for real-time edge synthesis via optimized latent diffusion.arXiv preprint arXiv:2504.08451, 2025

fields

years

verdicts

representative citing papers

citing papers explorer