{"id":"969fccf4-5140-42a2-b9a3-f500a8a6cc1e","arxiv_id":"2412.20496","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SGD weight-matrix eigenvalue fluctuations follow random matrix predictions, with variance proportional to learning rate divided by batch size, the linear scaling rule.","lead":"The paper uses random matrix theory, specifically Dyson Brownian motion, to describe how the eigenvalues of neural network weight matrices change during stochastic gradient descent. It derives the linear scaling rule between learning rate and batch size and tests it on a Gaussian restricted Boltzmann machine and a simple linear network.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unverified DBM noise isotropy in Eq. (5) is the load-bearing gap: SGD noise on W does not generically give the isotropic eigenbasis noise needed for the Coulomb-gas stationary distribution, and the linear-network test sidesteps this by injecting artificial isotropic noise.","rationale":"I considered two candidate concerns: (i) the finite-learning-rate/continuous-time limit, which the authors themselves flag in Sec. 2, and (ii) the validity of the eigenvalue SDE (5) itself. The reader's weakest_assumption was (i). I believe (ii) is more load-bearing because it precedes the continuous-time step: if Eq. (5) is not a correct description of the discrete eigenvalue process, no adjustment of the continuous-time limit can rescue the stationary Coulomb gas distribution. The paper gives no derivation of Eq. (5), only a reference to Ref. [2], and the stated relation between diagonal and off-diagonal noise variances is precisely the isotropy condition required for the DBM reduction. For SGD on a weight matrix W, the noise entering X = W^T W is multiplicative and depends on the instantaneous singular vectors; it is not generically isotropic. The RBM is a special solvable case and may satisfy the condition; the linear-network experiment intentionally adds isotropic noise, so it does not test the SGD-specific reduction. The RBM's successful Wigner fits and alpha/|B| scaling are real evidence for that model, but they do not establish the general claim. This supports the reader's CONDITIONAL verdict rather than changing it: the condition should explicitly include either a derivation of Eq. (5) from the SGD update (4) under transparent hypotheses, or direct empirical validation of the eigenbasis noise isotropy for a non-trivial SGD-trained network. I found no grounds to question the authors' integrity or the RBM numerics; the issue is purely a gap in the argument's generality.","tokens_in":10065,"tokens_out":11769,"duration_ms":118650,"concrete_test":"For the linear one-hidden-layer teacher-student setup of Sec. 4.2, train with genuine mini-batch SGD on the loss (13), without the ad hoc noise in Eq. (28). At several training times, record batch gradients and form the empirical covariance of \\delta X = (W+\\delta W)^T(W+\\delta W) - W^T W in the instantaneous eigenbasis of X. Estimate the diagonal variance C_{ii} and off-diagonal variance C_{ij} (i\\neq j) of the noise in that basis. Check the DBM condition 2 C_{ij} = C_{ii} (equivalently 2 \\tilde g_i^2 = V_B[\\Delta p]_{ii} with V_B[\\Delta p]_{i\\neq j} = \\tilde g_i^2). If the ratio differs substantially from 2, or if cross-correlations between different eigenbasis entries are non-negligible, Eq. (5) is not the correct eigenvalue SDE for SGD and the derived width scaling (9) lacks support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Eq. (5), the eigenvalue update for X = W^T W, quoted from Ref. [2] and not derived here. For a symmetric matrix process, the Ito reduction to Dyson Brownian motion produces a Coulomb term sum_j g_ij^2/(x_i - x_j) only when the noise in the instantaneous eigenbasis is isotropic: diagonal variance 2 g_i^2 and off-diagonal variance g_i^2, i.e., the relation 2\\tilde g_i^2 = V_B[\\Delta p]_{ii} = 2 V_B[\\Delta p]_{i\\neq j}. SGD noise enters W directly; the induced noise in X = W^T W is W^T \\delta W + \\delta W^T W plus a second-order term, and its covariance in the eigenbasis of X depends on the singular vectors of W. There is no general reason for this covariance to be isotropic or to satisfy the stated relation. If it is not, additional non-Coulomb noise-induced drift appears and the stationary measure is not the Coulomb gas (7). The two applications do not close the gap: the Gaussian RBM may possess the required special noise structure, and the linear teacher-student network (Sec. 4.2) is tested by adding ad hoc isotropic noise \\eta in Eq. (28) because the authors state the algorithm 'by itself is not noisy enough' — so the test verifies DBM with synthetic isotropic noise, not the reduction of SGD noise to DBM. Without an independent derivation or direct measurement of the noise covariance in the eigenbasis, the general linear-scaling rule (9) is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Park et al. present a lattice-field-theory proceedings contribution that applies Dyson Brownian motion and Coulomb-gas random matrix theory to the eigenvalue dynamics of weight matrices trained by stochastic gradient descent. From a central-limit representation of minibatch gradients (Eq. 4), they quote an eigenvalue Langevin equation (Eq. 5) with a Coulomb repulsion term, solve the associated Fokker-Planck equation for a stationary Coulomb gas (Eqs. 6-7), and obtain the variance formula sigma_i^2 = (alpha/|B|) V_B[Delta p]_{ii}/(2 Omega_i) (Eq. 9), which expresses the linear scaling rule that alpha/|B| fixes the fluctuation level. They test the resulting spectral density and Binder cumulant in a Gaussian RBM and introduce a linear one-hidden-layer teacher-student network in which an added noise term yields a two-species Coulomb-gas spectral density (Eq. 32).","tokens_in":10443,"tokens_out":15714,"duration_ms":154252,"significance":"If Eq. 9 is valid, the paper offers a physically clear separation of an optimiser-universal factor alpha/|B| from a model-dependent factor V_B/Omega, and it would provide a first-principles route to a widely used empirical scaling rule. The Gaussian-RBM evidence is the strongest part: the fitted spectral shape, the Binder cumulant near -0.147, and the independent variation of alpha and |B| in Fig. 3 support the predicted scaling form, and the data/code release in Ref. [26] makes the check reproducible. The value of the manuscript as submitted is nevertheless limited because the key reduction from SGD noise to Dyson Brownian motion is assumed rather than demonstrated, and the new linear-network test does not validate that reduction because it injects isotropic noise by hand (Eq. 28).","major_comments":[{"comment":"The derivation of the eigenvalue process is the load-bearing step, but Eq. (5) is quoted from Ref. [2] without proving that the DBM structure applies to SGD noise. For X = W^T W, the covariance of the noise in the eigenbasis is the image of the gradient-noise covariance under the singular-vector frame, and it is not isotropic for a generic loss. The DBM stationary measure (7) requires the special relation 2 tilde g_i^2 = V_B[Delta p]_{ii} = 2 V_B[Delta p]_{i neq j}: the noise variance must satisfy a precise relation between diagonal and off-diagonal entries in every instantaneous eigenbasis. This condition is not derived, and no simulation reported here measures it directly. Without it, additional noise-induced drift appears and Eq. (9) is not the stationary width. Please either supply the argument that SGD gradient covariance automatically satisfies this isotropy (which may be true for the Gaussian RBM class) or state explicitly that the derivation applies only to that restricted class.","section":"§3, Eq. (5)"},{"comment":"The linear teacher-student test is not evidence for the SGD-to-DBM reduction. Eq. (16) replaces the batch correlation by delta_{ij}, explicitly discarding the mini-batch stochasticity that is the source of SGD noise, and Eq. (28) then inserts an independent Gaussian eta with variance 0.01 because, in the authors' words, 'the algorithm by itself is not noisy enough.' Under this artificial noise prescription, the two-species spectral density (32) is expected by construction; the agreement in Fig. 5 therefore verifies the calculation of a two-species DBM fit, not the claim that SGD on a linear network produces that distribution. The abstract's claim that the paper derives the linear scaling rule for SGD remains unsupported by this section.","section":"§4.2, Eqs. (16), (28)"},{"comment":"The stationary calculation in Eqs. (6)-(7) is performed with a continuous-time Fokker-Planck equation, but the algorithm is the discrete map (1). The paper itself cites Refs. [17-19] for the failure of the naive alpha -> 0 limit to give a correct Ito stochastic differential equation. Since Eq. (5) is a discrete update, one needs a controlled stochastic-modified-equation limit to justify that the continuous-time drift and diffusion are those used in Eq. (6). Unless finite-learning-rate corrections are shown to be irrelevant for the stationary measure, the predicted Wigner semicircle and the alpha/|B| scaling are not guaranteed for discrete SGD. Please address this point explicitly.","section":"§2, transition to continuous time"},{"comment":"There is an inconsistency in the potential expansion. Eq. (8) defines tilde V_i = tilde V_i(x_s) + (1/(2 Omega_i))(x_i - x_s)^2 and calls Omega_i the curvature; with that convention, substituting Eq. (8) into the exponent of Eq. (7) gives sigma_i^2 = (alpha/|B|) Omega_i tilde g_i^2, not the tilde g_i^2/Omega_i form written in Eq. (9). If Eq. (9) is the intended result, Eq. (8) should read (Omega_i/2)(x_i - x_s)^2. Please correct the factor and define Omega_i unambiguously in terms of V''.","section":"§3, Eqs. (8) and (9)"}],"minor_comments":[{"comment":"The axis labels of the right panel are garbled in the arXiv text ('sqrt(alpha) |B| kappa^2 Omega'); please redraw the panel with clearly separated symbolic axis labels.","section":"Fig. 3 (right)"},{"comment":"Please state explicitly how many independent RBM runs and how many samples per run are used for the histograms and for the Binder cumulant, and report the measured U4 value with a statistical uncertainty rather than only showing the value in the figure.","section":"§4.1"},{"comment":"Because the paper summarizes Ref. [2] and adds a new section, it would be helpful to state explicitly which results are new to this contribution and which are taken from Ref. [2].","section":"Introduction and Ref. [2]"},{"comment":"The linear-network simulations are restricted to 2 x 2 matrices; while acceptable as a proof of concept, the claim of a 'generalised Wigner semi-circle' would be considerably strengthened by showing at least one larger-N example.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"This is a proceedings manuscript whose theoretical core is a summary of Ref. [2]; the only substantially new numerical section (Sec. 4.2) does not test the claimed SGD-to-DBM reduction. I would ask the authors to add either a direct measurement of the gradient-noise covariance in the eigenvalue frame or a clear restriction of the central claim to the Gaussian-RBM class, and to state explicitly how the present results go beyond Refs. [2] and [16]."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper is an honest conference proceedings that summarizes the authors' earlier DBM-based derivation of the linear scaling rule and adds one genuinely new result: a two-species Coulomb gas and its generalized Wigner semicircle (Eq. 32) for a linear one-hidden-layer network. The RBM validation is the strongest part of the paper. The eigenvalue histograms fit the Wigner semicircle, the Binder cumulant matches the predicted value, and the width scaling with sqrt(alpha/|B|) is tested by varying alpha and batch size independently. That is clean, reproducible evidence that the framework works for the Gaussian RBM.\n\nThe new linear-network result is a legitimate small contribution. The generalized spectral density in Eq. (32) naturally extends the Wigner semicircle to particles with different variances, and the fits in Fig. 5 show it beats the single-species form on peak and tails.\n\nThe soft spot, which the stress-test note correctly identifies, is load-bearing. Equation (5), the eigenvalue update, is quoted from Ref. [2] and not derived here. It requires the noise in the instantaneous eigenbasis to be isotropic: specifically, the diagonal variance 2 g~_i^2 and off-diagonal variance g~_i^2, tied to V_B[Delta p]_ii and V_B[Delta p]_i neq j. SGD noise acts on W, not directly on the eigenbasis of X = W^T W, and there is no general reason the induced covariance in that eigenbasis is isotropic. Without that, the stationary distribution is not the Coulomb gas (7), and the linear scaling rule (9) isn't established for arbitrary networks. The paper's tests don't close the gap. The RBM may have the needed special noise structure, but that's not shown either. And the linear-network test adds ad hoc isotropic noise (Eq. 28) because the algorithm 'by itself is not noisy enough,' so it validates DBM with synthetic noise, not the reduction of SGD noise to DBM.\n\nTo be fair, the authors are transparent about the continuous-time limit issue and about the injected noise. This is not a hidden flaw. It's an unproven assumption, and the paper's own empirical checks on a linear network deliberately bypass it. So the right framing is: conditional support for the DBM picture in a specific model, with a plausible but unproven generalization.\n\nWho's it for? People working on RMT for ML or on hyperparameter scaling, especially those who want the multi-species result. It's a proceedings paper, not a big journal submission, but it deserves a serious referee. My recommendation: send it to review. The referee should ask for a direct measurement or derivation of the eigenbasis noise covariance in SGD before the general linear scaling rule is stated, and should treat the two-species Coulomb gas as the main novel contribution.","headline":"Clean RBM evidence and a genuinely new two-species Coulomb gas result, but the general linear scaling rule rests on an unproven noise-isotropy assumption.","tokens_in":10973,"tokens_out":3728,"would_cite":true,"duration_ms":33149,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In the stationary limit, SGD-driven weight eigenvalues obey random matrix theory, with fluctuation width set by learning rate over batch size.","keywords":["stochastic gradient descent","random matrix theory","Dyson Brownian motion","Coulomb gas","linear scaling rule","Wigner semicircle","Restricted Boltzmann Machine","teacher-student network"],"falsifier":"Train a Gaussian RBM on a known target spectrum twice with the same $\\alpha/|B|$ but very different absolute values, for instance $\\alpha=0.01$, $|B|=10$ and $\\alpha=0.001$, $|B|=1$, and measure the width of an eigenvalue peak after convergence. If the two widths disagree beyond statistical error, or if the Binder cumulant of a doubly degenerate peak departs from $-4/27$, the universal scaling and Wigner spectral shape claimed by the paper are ruled out.","tokens_in":9872,"feed_emoji":"📊","tokens_out":7491,"duration_ms":67469,"temperature":0.7,"pith_summary":"The paper argues that the eigenvalues of weight matrices trained by stochastic gradient descent behave as a Coulomb gas in the stationary limit, so their statistics follow random matrix theory predictions such as the Wigner semicircle. The central quantitative claim is the variance formula $\\sigma_i^2 = \\frac{\\alpha}{|B|} \\frac{V_B[\\Delta p]_{ii}}{2\\Omega_i}$, which separates the optimizer-controlled factor $\\alpha/|B|$ from model-dependent gradient fluctuations. This derives the linear scaling rule—keeping the ratio of learning rate to batch size fixed keeps the fluctuation level of learning unchanged—from first-principles matrix dynamics. The claim is tested in a Gaussian Restricted Boltzmann Machine, where the fit to the Wigner semicircle and the Binder cumulant agree, and in a linear one-hidden-layer teacher-student network, where the spectral density becomes a generalized Wigner semicircle. If correct, the paper provides a physics origin for a widely used empirical training rule and a way to separate universal optimizer effects from architecture-specific ones.","feed_headline":"Eigenvalue noise in SGD scales with learning rate over batch size","feed_subtitle":"A random-matrix physics derivation explains a widely used rule for tuning neural network training.","key_machinery":"The central object is the eigenvalue Langevin equation obtained from Dyson Brownian motion for SGD, Eq. (5): $$x_i' = x_i + \\$\\alpha$ \\tilde K_i + \\frac{\\$alpha^{2}$}{|B|} \\sum_{j\\neq i} \\frac{\\tilde $g_i^{2}$}{x_i-x_j} + \\frac{\\$\\alpha$}{\\sqrt{|B|}} \\sqrt{2\\tilde $g_i^{2}$}\\,\\eta_i.$$ The Vandermonde determinant from the change of variables to eigenvalues produces the pairwise $1/(x_i-x_j)$ repulsion, which is what converts independent noise into Wigner statistics. The associated Fokker-Planck equation has a stationary Coulomb gas solution whose potential, expanded quadratically around each target eigenvalue, yields the variance formula Eq. (9).","core_discovery":"The paper establishes that in the stationary limit the eigenvalue distribution of weight matrices under SGD is governed by Dyson Brownian motion: each eigenvalue $x_i$ performs a drift-plus-noise motion with an additional Coulomb repulsion from every other eigenvalue, $1/(x_i-x_j)$. Solving the stationary Fokker-Planck equation gives a Coulomb gas distribution in which each eigenvalue fluctuates around its target value with a width $\\sigma_i^2 = \\frac{\\alpha}{|B|} \\frac{V_B[\\Delta p]_{ii}}{2\\Omega_i}$. The ratio $\\alpha/|B|$ enters solely through the stochasticity of the optimizer, while all model dependence sits in the ratio of gradient variance to potential curvature. The resulting spectral density around a learned eigenvalue is a Wigner semicircle, and nearest-neighbour spacings follow the Wigner surmise, as verified in the Gaussian RBM; with an extra linear layer the Coulomb gas acquires two species with different variances and the spectral density becomes a generalized Wigner semicircle while the level-spacing law survives.","pith_inferences":["If the stationary Coulomb gas description extends to finite learning rates in deeper networks, then the same ratio $\\alpha/|B|$ should govern the fluctuation level of layer-wise singular vectors and gradient statistics, not just eigenvalues.","A testable extension beyond the paper is to measure the eigenvalue distribution of a nonlinear teacher-student network with known target spectrum; the deviation of the fitted width from Eq. (9) would quantify where the continuous-time Langevin approximation breaks down.","The two-component Coulomb gas result suggests that in heterogeneous architectures, eigenvalue fluctuations should be characterized by a variance profile rather than a single scale, which could be used to detect layer-dependent training instabilities before they appear in the loss."],"forward_implications":["Scaling the learning rate and batch size together by the same factor leaves the stationary eigenvalue fluctuation width unchanged, giving a first-principles derivation of the practical linear scaling rule.","The spectral density around each learned eigenvalue is a Wigner semicircle, not a Gaussian, so the Binder cumulant $-4/27$ can be used as a model-independent signature of the Coulomb gas regime.","Adding hidden layers changes the effective potential curvature per eigenvalue, turning the Coulomb gas into a multi-species system; the level-spacing Wigner surmise survives, but the spectral density generalizes and develops wider tails.","Fluctuation control separates cleanly into hyperparameters ($\\alpha/|B|$) and model properties ($V_B[\\Delta p]_{ii}/(2\\Omega_i)$), so each can be tuned or measured independently."],"supporting_citations":[{"why":"Supplies the eigenvalue Langevin equation Eq. (5), the Coulomb repulsion term, and the first derivation of the linear scaling rule from matrix dynamics.","marker":"[2]"},{"why":"Dyson Brownian motion; provides the framework that gives the eigenvalue process and the Coulomb gas stationary distribution.","marker":"[8]"},{"why":"Mehta's Random Matrices; source for Coulomb gas methods, Wigner semicircle, and Wigner surmise used to make predictions.","marker":"[9]"},{"why":"Reports the empirical linear scaling rule in large-batch ImageNet training that this paper derives from first principles.","marker":"[12]"},{"why":"Warns that the naive alpha to 0 continuous-time limit is not a correct stochastic differential equation, motivating the treatment used here.","marker":"[17]"},{"why":"Provides fluctuation-dissipation relations for SGD, another reference for the continuous-time limit issue.","marker":"[19]"},{"why":"Gives the Gaussian RBM analysis and lattice-field target spectrum used to test the Coulomb gas predictions.","marker":"[23]"}],"fun_headline_variants":["SGD weight dynamics follow Dyson Brownian motion","Random matrix theory explains SGD learning rate rule","Eigenvalue repulsion shapes neural net weight spectra","Dyson Brownian motion describes SGD weight matrices","Universal spectral law for stochastic gradient descent"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole derivation hinges on the assumption that discrete SGD can be represented by a continuous-time Langevin and Fokker-Planck process whose stationary solution is the Coulomb gas; if finite-learning-rate corrections change that stationary distribution, the predicted Wigner semicircle and $\\sqrt{\\alpha/|B|}$ width scaling would not hold exactly for actual discrete updates.","fun_headline_variants_meta":{"raw":{"variants":["SGD weight dynamics follow Dyson Brownian motion","Random matrix theory explains SGD learning rate rule","Eigenvalue repulsion shapes neural net weight spectra","Dyson Brownian motion describes SGD weight matrices","Universal spectral law for stochastic gradient descent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2722,"prompt_tokens":842,"completion_tokens":1880,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":1810}},"tokens_in":458,"tokens_out":1880,"duration_ms":12443,"temperature":1.0,"reasoning_tokens":1810,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:20:33.719147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a Gaussian RBM on a known target spectrum twice with the same $\\alpha/|B|$ but very different absolute values, for instance $\\alpha=0.01$, $|B|=10$ and $\\alpha=0.001$, $|B|=1$, and measure the width of an eigenvalue peak after convergence. If the two widths disagree beyond statistical error, or if the Binder cumulant of a doubly degenerate peak departs from $-4/27$, the universal scaling and Wigner spectral shape claimed by the paper are ruled out.","supporting_citations":[{"cited_title":"Dyson,A Brownian-Motion Model for the Eigenvalues of a Random Matrix, J","cited_arxiv_id":null,"evidence_quote":"Dyson Brownian motion; provides the framework that gives the eigenvalue process and the Coulomb gas stationary distribution."},{"cited_title":"Mehta,Random Matrices, Academic Press, New York, 3rd ed","cited_arxiv_id":null,"evidence_quote":"Mehta's Random Matrices; source for Coulomb gas methods, Wigner semicircle, and Wigner surmise used to make predictions."},{"cited_title":"Mandt, M.D","cited_arxiv_id":null,"evidence_quote":"Warns that the naive alpha to 0 continuous-time limit is not a correct stochastic differential equation, motivating the treatment used here."}],"review_version":1}