Pith. sign in

REVIEW 1 cited by

Dyson Brownian motion and random matrix dynamics of weight matrices during learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.13512 v1 pith:GZESF6OG submitted 2024-11-20 cond-mat.dis-nn cs.LGhep-lat

classification cond-mat.dis-nncs.LGhep-lat
keywords dynamicslearningmatrixweightbrownianduringdysonlinear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

During training, weight matrices in machine learning architectures are updated using stochastic gradient descent or variations thereof. In this contribution we employ concepts of random matrix theory to analyse the resulting stochastic matrix dynamics. We first demonstrate that the dynamics can generically be described using Dyson Brownian motion, leading to e.g. eigenvalue repulsion. The level of stochasticity is shown to depend on the ratio of the learning rate and the mini-batch size, explaining the empirically observed linear scaling rule. We verify this linear scaling in the restricted Boltzmann machine. Subsequently we study weight matrix dynamics in transformers (a nano-GPT), following the evolution from a Marchenko-Pastur distribution for eigenvalues at initialisation to a combination with additional structure at the end of learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Random Matrix Theory for Stochastic Gradient Descent

    hep-lat 2024-12 conditional novelty 4.0 of 10

    SGD weight-matrix eigenvalue fluctuations follow random matrix predictions, with variance proportional to learning rate divided by batch size, the linear scaling rule.

Pith tools