Pith. sign in

Dyson Brownian motion and random matrix dynamics of weight matrices during learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

During training, weight matrices in machine learning architectures are updated using stochastic gradient descent or variations thereof. In this contribution we employ concepts of random matrix theory to analyse the resulting stochastic matrix dynamics. We first demonstrate that the dynamics can generically be described using Dyson Brownian motion, leading to e.g. eigenvalue repulsion. The level of stochasticity is shown to depend on the ratio of the learning rate and the mini-batch size, explaining the empirically observed linear scaling rule. We verify this linear scaling in the restricted Boltzmann machine. Subsequently we study weight matrix dynamics in transformers (a nano-GPT), following the evolution from a Marchenko-Pastur distribution for eigenvalues at initialisation to a combination with additional structure at the end of learning.

citation-role summary

background 1

citation-polarity summary

fields

hep-lat 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Random Matrix Theory for Stochastic Gradient Descent

hep-lat · 2024-12-29 · conditional · novelty 4.0

SGD weight-matrix eigenvalue fluctuations follow random matrix predictions, with variance proportional to learning rate divided by batch size, the linear scaling rule.

citing papers explorer

Showing 1 of 1 citing paper.

  • Random Matrix Theory for Stochastic Gradient Descent hep-lat · 2024-12-29 · conditional · none · ref 16 · internal anchor

    SGD weight-matrix eigenvalue fluctuations follow random matrix predictions, with variance proportional to learning rate divided by batch size, the linear scaling rule.