Aurora is a leverage-aware spectral optimizer that enforces uniform row norms in matrix updates while preserving Muon's polar geometry, outperforming Muon and achieving SOTA among spectral methods on modded-nanoGPT.
Dying ReLU and initialization: Theory and numerical examples
4 Pith papers cite this work, alongside 248 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 4roles
background 1polarities
background 1representative citing papers
Plasticity loss in GPT-style transformers on multilingual tasks persists from 5M to 314M parameters, follows a sublinear scaling law with model size, and occurs in both continual and stationary settings.
Bilevel-optimized implicit neural representation with Gaussian process hyperparameter tuning enables scan-specific accelerated MRI reconstruction without training data.
Multi-narrow single-model ensembles outperform wide baselines in low-data image classification by learning diverse features but underperform in data-rich settings where training favors few paths.
citing papers explorer
-
Aurora: A Leverage-Aware Spectral Optimizer
Aurora is a leverage-aware spectral optimizer that enforces uniform row norms in matrix updates while preserving Muon's polar geometry, outperforming Muon and achieving SOTA among spectral methods on modded-nanoGPT.
-
Can Scale Save Us From Plasticity Loss in Large Language Models?
Plasticity loss in GPT-style transformers on multilingual tasks persists from 5M to 314M parameters, follows a sublinear scaling law with model size, and occurs in both continual and stationary settings.
-
Bilevel Optimized Implicit Neural Representation for Scan-Specific Accelerated MRI Reconstruction
Bilevel-optimized implicit neural representation with Gaussian process hyperparameter tuning enables scan-specific accelerated MRI reconstruction without training data.
-
Multi-Narrow Transformation as a Single-Model Ensemble: Boundary Conditions, Mechanisms, and Failure Modes
Multi-narrow single-model ensembles outperform wide baselines in low-data image classification by learning diverse features but underperform in data-rich settings where training favors few paths.