REVIEW 4 major objections 6 minor 1 cited by
Absorbing state dynamics of stochastic gradient descent
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Stochastic gradient descent reduces to biased random organization near a critical packing fraction of about 0.64.
desk verdict A clean new mapping between SGD and BRO, with solid numerics, but the Gaussian substitution is the weak point in the analytic equivalence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the mean-covariance (Gaussian) approximation of the discrete update noise: each random kick (BRO) or random batch selection (SGD) is replaced by a Gaussian with matching first and second moments, turning both processes into stochastic differential equations with anisotropic multiplicative noise. For pairwise BRO this is $dx_i = -\frac{\epsilon}{\tau}\nabla_i V\,dt + \frac{\epsilon}{\sqrt{3\tau}}\sum_j\sqrt{\Lambda_{ji}}\,dW_{ji}$, and for SGD it is the analogous equation with prefactors $\alpha b_f$ for the drift and $\alpha\sqrt{b_f-b_f^2}$ for the noise. The two SDEs match exactly under the parameter map $\alpha=\epsilon/b_f$, $b_f=3/4$, so the approximation is the load-bearing identity that makes the dynamical equivalence a theorem about the two original algorithms.
What would settle it
Compare the critical packing fraction and the steady-state activity scaling of SGD and BRO at a fixed small $\alpha=\epsilon/b_f$, $b_f=3/4$, but replace the Bernoulli batch selection with a heavy-tailed or skewed selection rule that keeps the same mean and covariance; if $\phi_c$ or the measured exponents move away from $0.64$ and the Manna values, the higher noise cumulants are relevant and the claimed equivalence would fail.
Extended reading notes
Core claim
Starting from the pairwise BRO update, in which each overlapping pair is displaced by equal-magnitude random kicks along their line of centers, the authors decompose each kick into its mean and its fluctuation and replace the latter by Gaussian noise with the same covariance. The mean is the negative gradient of a linear repulsive potential $U(r)=\frac12(2R-|r|)$, and the fluctuation becomes an anisotropic multiplicative noise term $\frac{\epsilon}{\sqrt{3}}\sum_j \sqrt{\Lambda_{ji}}\,\xi_{ji}$. Applying the same mean-covariance decomposition to an energy-based SGD that selects active pairs with probability $b_f$ and moves them by $\alpha$-sized gradient steps yields exactly the same stochastic approximation when $\alpha=\epsilon/b_f$ and $b_f=3/4$. The paper therefore establishes that the two dynamics are equivalent in distribution at small kick sizes, and verifies numerically that their mean square displacements, structure factors, and finite-size scaling agree. At the critical point both schemes reach the random-close-packing fraction $\phi_c\approx0.64$ as $\epsilon,\alpha\to0$, and the activity and relaxation-time exponents match the Manna universality class values ($\beta=0.84$, $\nu_\parallel=1.08$) regardless of batch size.
Load-bearing premise
The Gaussian mean-covariance approximation of the discrete update noise is assumed to preserve the absorbing-state critical behavior, even though only the first two noise moments are matched and the paper validates the approximation numerically at small kick sizes rather than proving that higher cumulants are irrelevant.
Editorial extensions
If this is right
- As the learning rate goes to zero, SGD's critical packing fraction rises toward the random-close-packing value $\phi_c \approx 0.64$, regardless of batch fraction.
- Near the transition, the steady-state activity and relaxation time of SGD and RCD collapse onto Manna-universality scaling curves, including at $b_f=1$ where SGD reduces to gradient descent.
- Because the stochastic approximation matches exactly, the equivalence implies that the noise statistics of minibatch selection and BRO's kick-size statistics play identical roles in the critical region.
- Above the critical point, SGD with small batch sizes settles into flatter minima than gradient descent, whereas RCD settles into sharper ones.
Reading between the lines
- Because only the mean and covariance of the noise are matched, the equivalence suggests that higher-order noise statistics (skewness, heavy tails) are irrelevant for the absorbing transition; the paper does not prove this, and it is a testable prediction of the stochastic-approximation framework.
- The explicit SDE approximations allow the authors' conjecture of mean-field behavior for $d\ge4$ to be tested in simulation without running the full discrete algorithms, since integrating the SDEs at $d=4$ and $d=5$ would show whether the Manna exponents cross over.
- In practical representation learning, the result implies that batch size and learning rate are physical control parameters: they place the learning dynamics on one side or the other of an absorbing transition between fully separated and partially overlapping neural manifolds, a perspective the paper opens but does not develop into algorithmic advice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies a minimal model of 'neural manifold packing' in which stochastic gradient descent (SGD) is applied to spherical particles with short-ranged linear repulsive interactions, and compares it with biased random organization (BRO), a nonequilibrium absorbing-state model. The authors derive Gaussian approximations to pairwise BRO and to their energy-based SGD by matching the mean and covariance of the discrete updates. They show that, at learning rate α=ε/b_f and batch fraction b_f=3/4, the two Gaussian-approximated processes coincide (Eqs. 4 and 6), and similarly for individual-particle BRO and random coordinate descent (RCD). Numerical simulations report mean-square displacement agreement at small kick sizes, critical packing fractions φc→0.64 as ε,α→0, finite-size scaling consistent with the Manna universality class for both SGD and RCD at various b_f (including b_f=1, i.e., deterministic gradient descent), and a flatness bias for SGD above φc.
Significance. The conceptual bridge between a machine-learning optimization algorithm and absorbing-state critical phenomena is attractive. The mean-covariance computations in the Appendix are transparent, and I verified the exact parameter matching at α=ε/b_f, b_f=3/4; within the Gaussian approximation, this is a genuine parameter-free derivation. The external benchmark against the known BRO value φc≈0.64 is a nice sanity check. However, the paper's central claims rest on the assumption that replacing the true discrete noises by Gaussians does not alter the absorbing-state critical behavior, and this assumption is currently tested only by MSD agreement away from criticality. The Manna-universality claim is also weakened by the a priori fixing of exponents in the finite-size scaling analysis. If the proposed near-critical tests are passed, the paper would make a valuable contribution to the statistical-mechanics understanding of SGD-like dynamics.
major comments (4)
- [Stochastic approximations, Eqs. (2) and (6)] The central equivalence between BRO and SGD is derived only after replacing the true discrete noises by Gaussian noises with matched mean and covariance; Eq. (6) equals Eq. (4) exactly only for these Gaussian-approximated processes. The validation in Fig. 5(a) and Fig. S1 compares mean-square displacements at φ=0.63 for small kick sizes, which does not probe the absorbing-state transition. Near φc activity is sparse and each particle receives very few kicks before freezing, so no central-limit argument justifies the replacement; the higher cumulants of the uniform (BRO) and Bernoulli (SGD) noises, which differ, could in principle change φc or the universality class. A concrete test would be to compare the exact BRO/SGD dynamics with their Gaussian approximations for the steady-state activity, survival probability, and finite-size scaling near φc at the same small ε; without such a test the claimed equivalence is conditional on an unverified assumption.
- [Critical behavior, Fig. 3 and Manna universality] The claim that SGD and RCD belong to the Manna universality class is supported by finite-size scaling analyses in which the Manna exponents β=0.84, ν∥=1.08, and ν⊥=0.59 are fixed a priori and only φc is fitted; the text states that this is done 'assuming the Manna exponents'. This is a consistency check, not a measurement, because imposing the exponents cannot distinguish Manna from nearby universality classes. A stronger test would fit β and ν∥ as free parameters, or at least compare the collapse quality against alternative exponent sets. The paper should either soften the claim to 'consistent with Manna exponents' or provide such a free-exponent analysis.
- [Critical behavior, b_f=1.0 paragraph] For b_f=1.0, SGD and RCD reduce to deterministic gradient descent, which the paper itself notes is not an absorbing-state model because it 'absorbs' on both sides of φc. The finite-size scaling shown in Fig. 3(c,d) for b_f=1.0 is therefore not obviously well-defined: for φ>φc the steady state is a static minimum, not an active fluctuating steady state, and the paper does not specify how fa∞ and τr are computed in that case. The claim that the zero-noise limit is also in the Manna universality class needs either a precise definition of the observables for GD or a separate argument that the deterministic limit is a well-defined absorbing-state dynamics.
- [Flatness analysis, Fig. 4 and Eq. (S12)] The theoretical justification of the flatness measure, ΔV ∝ Tr(H), assumes a second-order smooth potential, but the model potential U(r)=1/2(2R−r) for r<2R is piecewise linear: its Hessian is zero in the interior of the overlap region and singular at the contact point r=2R. The relationship ΔV ∝ Tr(H) therefore does not follow for this model, and contributions from the potential's kink may dominate the energy fluctuation. The authors should either use a smooth potential for the flatness comparison or justify the flatness measure directly without the Hessian expansion. In addition, Fig. 4(b) shows no error bars for the reported energy fluctuations.
minor comments (6)
- [Appendix A, final paragraph] The sentence 'the stochastic approximations of individual-particle BRO and SGD are equivalent' should read 'individual-particle BRO and RCD', since Eq. (A8) is matched to the individual-particle BRO approximation, while SGD is matched to pairwise BRO in Eq. (6).
- [Algorithm 2, Supplementary Material] The active-particle condition contains '|xj_k−xj_k|<2R'; this should be '|xj_k−xi_k|<2R'.
- [Supplementary Material, Eq. (S5)] Equation (S5) is missing a minus sign on the drift term; it should read dx(t) = −ε/τ ∇iV dt + ..., consistent with Eq. (5) in the main text.
- [Methods, definition of b_f] The text defines b_f as the 'fraction of pairs in a batch', but Algorithm 3 selects each active pair independently with probability b_f, so the realized batch fraction is random; please clarify whether b_f denotes the expected fraction or a fixed batch size.
- [Fig. 5(b)] The convergence of φc to φRCP as α→0 is presented without error bars or a specified fitting procedure for the extrapolation; reporting the fitted values and their statistical uncertainties would strengthen the claim.
- [General] The paper would benefit from a data/code availability statement, given that all central results are numerical.
Assumptions & free parameters
free parameters (2)
- batch fraction bf at the exact-matching point =
3/4
- dimensionless perturbation scale for the flatness measure =
0.03
assumptions (5)
- domain assumption Mean-covariance Gaussian approximation preserves the dynamics of the discrete stochastic processes near the critical point.
- domain assumption Neural manifolds in a 3D embedded state space can be represented as hard spheres, and their separation during learning is a sphere-packing problem.
- domain assumption BRO belongs to the Manna universality class with exponents β=0.84, ν∥=1.08, ν⊥=0.59 and has critical density φc≈0.64.
- domain assumption The upper critical dimension of the Manna class is d=4.
- standard math Kabatyanskii-Levenshtein upper bound φv(d) ≲ 2^{-0.599d} for sphere packings in high dimensions.
Cite this review
Pith. "Pith review of Absorbing state dynamics of stochastic gradient descent." pith.science (2026). https://pith.science/paper/YSZGNMAU
@misc{pith2026241111834,
author = {Pith},
title = {Pith review of: Absorbing state dynamics of stochastic gradient descent},
year = {2026},
howpublished = {\url{https://pith.science/paper/YSZGNMAU}},
note = {Machine review of arXiv:2411.11834}
}
abstract
Stochastic gradient descent (SGD) is a fundamental tool for training deep neural networks across a variety of tasks. In self-supervised learning, different input categories map to distinct manifolds in the embedded neural state space. Accurate classification is achieved by separating these manifolds during learning, akin to a packing problem. We investigate the dynamics of ``neural manifold packing'' by employing a minimal model in which SGD is applied to spherical particles in physical space. In this model, SGD minimizes the system's energy (classification loss) by stochastically reducing overlaps between particles (manifolds). We observe that this process undergoes an absorbing phase transition, prompting us to use the framework of biased random organization (BRO), a nonequilibrium absorbing state model, to describe SGD behavior. We show that BRO dynamics can be approximated by those of particles with linear repulsive interactions under multiplicative anisotropic noise. Thus, for a linear repulsive potential and small kick sizes (learning rates), we find that BRO and SGD become equivalent, converging to the same critical packing fraction $\phi_c \approx 0.64$, despite the fundamentally different origins of their noise. This equivalence is further supported by the observation that, like BRO, near the critical point, SGD exhibits behavior consistent with the Manna universality class. Above the critical point, SGD exhibits a bias towards flatter minima of the energy landscape, reinforcing the analogy with loss minimization in neural networks.
Figures
Forward citations
Cited by 1 Pith paper
-
Hyperuniform systems are maximally irreversible
Across four noisy particle systems, entropy production rate peaks precisely at the hyperuniformity threshold, and a path-integral derivation shows this follows from the non-invertibility of the diffusion tensor at the...
Reference graph
Works this paper leans on
-
[1]
S. Bubeck et al. , Foundations and Trends in Machine Learning 8, 231 (2015)
work page 2015
-
[2]
C. M. Bishop and N. M. Nasrabadi,Pattern recognition and machine learning , Vol. 4 (Springer, 2006). 6
work page 2006
-
[3]
C. M. Bishop and H. Bishop,Deep learning : foundations and concepts (Springer, 2024)
work page 2024
-
[4]
Bengio, inNeural networks: Tricks of the trade: Sec- ond edition (Springer, 2012) pp
Y. Bengio, inNeural networks: Tricks of the trade: Sec- ond edition (Springer, 2012) pp. 437–478
work page 2012
-
[5]
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, Na- ture 323, 533 (1986)
1986
-
[6]
L. Bottou and Y. Cun, Advances in neural information processing systems 16 (2003)
work page 2003
-
[7]
Y. Bengio, inProceedings of ICML workshop on unsuper- vised and transfer learning (JMLR Workshop and Con- ference Proceedings, 2012) pp. 17–36
work page 2012
-
[8]
Q. Wang, Y. Ma, K. Zhao, and Y. Tian, Annals of Data Science , 1 (2020)
work page 2020
Show all 71 references
- [9]
-
[10]
Geiger, S
M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity- Jesi, G. Biroli, and M. Wyart, Physical Review E100, 012115 (2019)
2019
-
[11]
Pellegrini and G
F. Pellegrini and G. Biroli, Advances in Neural Informa- tion Processing Systems33, 5356 (2020)
2020
-
[12]
Krizhevsky, I
A. Krizhevsky, I. Sutskever, and G. E. Hinton, Ad- vances in neural information processing systems 25, 10.1145/3065386 (2012)
2012 doi
- [13]
-
[14]
Jacot, F
A. Jacot, F. Gabriel, and C. Hongler, Advances in neural information processing systems31 (2018)
2018
-
[15]
Gerbelot, E
C. Gerbelot, E. Troiani, F. Mignacco, F. Krzakala, and L. Zdeborova, SIAM Journal on Mathematics of Data Science 6, 400 (2024)
2024
-
[16]
Q. Li, C. Tai, and E. Weinan, inInternational Conference on Machine Learning (PMLR, 2017) pp. 2101–2110
2017
-
[18]
Rotskoff and E
G. Rotskoff and E. Vanden-Eijnden, Communications on Pure and Applied Mathematics75, 1889 (2022)
2022
-
[19]
N. Yang, C. Tang, and Y. Tu, Physical Review Letters 130, 237101 (2023)
2023
-
[20]
S. Mei, A. Montanari, and P.-M. Nguyen, Proceedings of the National Academy of Sciences115, E7665 (2018)
2018
-
[21]
Azizian, F
W. Azizian, F. Iutzeler, J. Malick, and P. Mertikopoulos, arXiv preprint arXiv:2406.09241 (2024)
2024 arXiv
-
[22]
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein, Advances in Neural Information Processing Systems31 (2018)
2018
-
[23]
S. K. Ainsworth, J. Hayase, and S. Srinivasa, arXiv preprint arXiv:2209.04836 (2022)
2022 arXiv
-
[24]
Stringer, M
C. Stringer, M. Pachitariu, N. Steinmetz, M. Carandini, and K. D. Harris, Nature571, 361 (2019)
2019
-
[25]
Stringer, M
C. Stringer, M. Pachitariu, N. Steinmetz, C. B. Reddy, M. Carandini, and K. D. Harris, Science364, eaav7893 (2019)
2019
-
[26]
Kafashan, A
M. Kafashan, A. W. Jaffe, S. N. Chettih, R. Nogueira, I. Arandia-Romero, C. D. Harvey, R. Moreno-Bote, and J. Drugowitsch, Nature Communications12, 473 (2021)
2021
-
[27]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, in International Conference on Machine Learning (PMLR,
-
[28]
9729–9738
K.He, H.Fan, Y.Wu, S.Xie,andR.Girshick,in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2020) pp. 9729–9738
2020
-
[29]
Goyal, Q
P. Goyal, Q. Duval, I. Seessel, M. Caron, I. Misra, L. Sagun, A. Joulin, and P. Bojanowski, arXiv preprint arXiv:2202.08360 (2022)
2022 arXiv
-
[30]
J. D. Bernal and J. Mason, Nature188, 910 (1960)
1960
-
[31]
R. P. Behringer and B. Chakraborty, Reports on Progress in Physics 82, 012601 (2018)
2018
-
[32]
J. H. Conway and N. J. A. Sloane,Sphere packings, lat- tices and groups , Vol. 290 (Springer Science & Business Media, 2013)
2013
-
[33]
Chung, D
S. Chung, D. D. Lee, and H. Sompolinsky, Physical Re- view X 8, 031003 (2018)
2018
-
[34]
Chung, D
S. Chung, D. D. Lee, and H. Sompolinsky, Physical Re- view E 93, 060301 (2016)
2016
-
[35]
Mignacco, C.-N
F. Mignacco, C.-N. Chou, and S. Chung, arXiv preprint arXiv:2405.06851 (2024)
2024 arXiv
-
[36]
Corte, P
L. Corte, P. M. Chaikin, J. P. Gollub, and D. J. Pine, Nature Physics 4, 420 (2008)
2008
-
[37]
Wilken, R
S. Wilken, R. E. Guerra, D. J. Pine, and P. M. Chaikin, Phys. Rev. Lett.125, 148001 (2020)
2020
-
[38]
Tjhung and L
E. Tjhung and L. Berthier, Physical review letters114, 148301 (2015)
2015
-
[39]
Wilken, R
S. Wilken, R. E. Guerra, D. Levine, and P. M. Chaikin, Physical Review Letters127, 038002 (2021)
2021
-
[40]
Wilken, A
S. Wilken, A. Z. Guo, D. Levine, and P. M. Chaikin, Phys. Rev. Lett.131, 238202 (2023)
2023
-
[41]
Galliano, M
L. Galliano, M. E. Cates, and L. Berthier, Phys. Rev. Lett. 131, 047101 (2023)
2023
-
[42]
Hexner and D
D. Hexner and D. Levine, Phys. Rev. Lett.118, 020601 (2017)
2017
-
[43]
Mignacco, F
F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová, Advances in Neural Information Processing Systems33, 9540 (2020)
2020
-
[44]
Mignacco and P
F. Mignacco and P. Urbani, Journal of Statistical Me- chanics: Theory and Experiment2022, 083405 (2022)
2022
-
[45]
R. D. Kamien and A. J. Liu, Physical review letters99, 155501 (2007)
2007
-
[46]
Donev, F
A. Donev, F. H. Stillinger, and S. Torquato, Physical review letters 95, 090604 (2005)
2005
-
[47]
K. W. Desmond and E. R. Weeks, Physical Review E80, 051305 (2009)
2009
-
[48]
Anzivino, M
C. Anzivino, M. Casiulis, T. Zhang, A. S. Moussa, S. Martiniani, and A. Zaccone, The Journal of Chemi- cal Physics 158 (2023)
2023
-
[49]
Ness and M
C. Ness and M. E. Cates, Physical review letters124, 088004 (2020)
2020
-
[50]
W. Hu, C. J. Li, L. Li, and J.-G. Liu, Annals of Mathe- matical Sciences and Applications4 (2019)
2019
-
[51]
Mandt, M
S. Mandt, M. D. Hoffman, and D. M. Blei, Journal of Machine Learning Research18, 1 (2017)
2017
-
[52]
Lennard and I
J. Lennard and I. Jones, Proceedings of the Royal Society of London. Series A, Containing Papers of a Mathemat- ical and Physical Character106, 441 (1924)
1924
-
[53]
J. D. Weeks, D. Chandler, and H. C. Andersen, The Jour- nal of Chemical Physics54, 5237 (1971)
1971
-
[54]
S. J. Wright, Mathematical programming151, 3 (2015)
2015
-
[55]
Martiniani, P
S. Martiniani, P. M. Chaikin, and D. Levine, Physical Review X 9, 011031 (2019)
2019
-
[56]
E.TjhungandL.Berthier,JournalofStatisticalMechan- ics: Theory and Experiment2016, 033501 (2016)
2016
-
[57]
Henkel, H
M. Henkel, H. Hinrichsen, and S. Lubeck, Non- Equilibrium Phase Transitions , 2009th ed., Theoretical and mathematical physics (Springer, New York, NY, 2009)
2009
-
[58]
Sorge, Zenodo: A scientific Python package for finite- size scaling analysis (2015)
A. Sorge, Zenodo: A scientific Python package for finite- size scaling analysis (2015)
2015
-
[59]
Jastrzębski, Z
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fis- cher, Y. Bengio, and A. Storkey, arXiv preprint 7 arXiv:1711.04623 (2017)
2017 arXiv
-
[60]
Z. Xie, I. Sato, and M. Sugiyama, arXiv preprint arXiv:2002.03495 (2020)
2020 arXiv
-
[61]
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyan- skiy, and P. T. P. Tang, arXiv preprint arXiv:1609.04836 (2016)
2016 arXiv
-
[62]
I Can ’t Believe It’s Not Better!
W. R. Huang, Z. Emam, M. Goldblum, L. Fowl, J. K. Terry, F. Huang, and T. Goldstein, in Proceedings on "I Can ’t Believe It’s Not Better!" at NeurIPS Work- shops, Proceedings of Machine Learning Research, Vol. 137, edited by J. Zosa Forde, F. Ruiz, M. F. Pradier, and A. Schein...
2020
-
[63]
T. C. Hales, arXiv preprint math/9811071 (1998)
1998 arXiv
-
[64]
Zaccone, Phys
A. Zaccone, Phys. Rev. Lett.128, 028002 (2022)
2022
-
[65]
Donev, I
A. Donev, I. Cisse, D. Sachs, E. A. Variano, F. H. Still- inger, R. Connelly, S. Torquato, and P. M. Chaikin, Sci- ence 303, 990 (2004)
2004
-
[66]
Rocks and R
S. Rocks and R. S. Hoy, Soft Matter19, 5701 (2023)
2023
-
[67]
S. Ai, D. Levine, and C. Paul, (Private Communication) (2025)
2025
-
[68]
G. A. Kabatiansky and V. I. Levenshtein, Problemy peredachi informatsii 14, 3 (1978)
1978
-
[69]
Cohn, arXiv preprint arXiv:1611.01685 (2016)
H. Cohn, arXiv preprint arXiv:1611.01685 (2016). Appendix Stochastic approximation of individual-particle BRO. In individual-particle BRO, the update rule for active particlei at iterationk reads, xi k+1 = xi k +ui k ∑ j∈Γi ∆xji k. (A1) where ∆xji k =− xj k−xi k |xj k−xi k|. u...
2016 arXiv
-
[70]
Q. Li, C. Tai, and E. Weinan, The Journal of Machine Learning Research 20, 1474 (2019)
2019
-
[71]
Mandt, M
S. Mandt, M. D. Hoffman, and D. M. Blei, Journal of Machine Learning Research 18, 1 (2017)
2017
-
[72]
Wilken, R
S. Wilken, R. E. Guerra, D. Levine, and P. M. Chaikin, Physical Review Letters 127, 038002 (2021)
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.