SGD weight-matrix eigenvalue fluctuations follow random matrix predictions, with variance proportional to learning rate divided by batch size, the linear scaling rule.
Dyson,Statistical theory of the energy levels of complex systems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
hep-lat 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Random Matrix Theory for Stochastic Gradient Descent
SGD weight-matrix eigenvalue fluctuations follow random matrix predictions, with variance proportional to learning rate divided by batch size, the linear scaling rule.