Effective noise scale non-monotonically governs model merging success with an optimum, unifying effects of learning rate, weight decay, batch size, and augmentation on the loss landscape.
Optimizers qualitatively alter solutions and we should leverage this
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 5roles
background 1polarities
background 1representative citing papers
Feedback alignment in deep networks is limited by low-rank error signals; orthogonal weight updates and activity normalization raise effective rank and boost performance.
The same Transformer architecture follows different spectral scaling laws under different optimizers, with Muon achieving linear hard-rank scaling on tail representations while AdamW shows weak scaling, even when perplexity is matched.
Muon optimizer improves performance over Adam in equivariant networks on ModelNet40 and produces solutions with larger Hessian curvature, more regular loss surfaces, and higher stable/effective ranks.
Shallow MLPs and dense CPGs outperform deeper MLPs and Actor-Critic RL in bounded robot control tasks with limited proprioception, with a Parameter Impact metric indicating extra RL parameters yield no performance gain over evolutionary strategies.
citing papers explorer
-
How does the optimizer implicitly bias the model merging loss landscape?
Effective noise scale non-monotonically governs model merging success with an optimum, unifying effects of learning rate, weight decay, batch size, and augmentation on the loss landscape.
-
Overcoming Rank Collapse in Feedback Alignment
Feedback alignment in deep networks is limited by low-rank error signals; orthogonal weight updates and activity normalization raise effective rank and boost performance.
-
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
The same Transformer architecture follows different spectral scaling laws under different optimizers, with Muon achieving linear hard-rank scaling on tail representations while AdamW shows weak scaling, even when perplexity is matched.
-
How the Optimizer Shapes Learned Solutions in Equivariant Neural Networks
Muon optimizer improves performance over Adam in equivariant networks on ModelNet40 and produces solutions with larger Hessian curvature, more regular loss surfaces, and higher stable/effective ranks.
-
Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization
Shallow MLPs and dense CPGs outperform deeper MLPs and Actor-Critic RL in bounded robot control tasks with limited proprioception, with a Parameter Impact metric indicating extra RL parameters yield no performance gain over evolutionary strategies.