{"id":"726ea072-76d7-4a52-8b77-1d63ee7968ac","arxiv_id":"2411.11834","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SGD on linearly repulsive particles is mapped onto biased random organization, sharing its critical packing fraction of about 0.64 and its Manna universality class.","lead":"Stochastic gradient descent, the standard training algorithm for neural networks, is shown to behave like a physical packing process in a minimal sphere model. In the limit of small learning rates it becomes equivalent to biased random organization, with both processes reaching the same critical density near 0.64.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Gaussian replacement of the kick and pair-selection noise in Eqs. 2 and 6 is the load-bearing step for the BRO–SGD equivalence, and it is checked only by MSD agreement, not by comparing the critical behavior of the exact and approximated processes.","rationale":"I read the paper as making two connected claims: first, that BRO and SGD become equivalent for a linear repulsive potential at small kick sizes, with Eq. 6 matching Eq. 4 under the parameter mapping; second, that both lie in the Manna universality class near the resulting critical point. The analytic matching of means and covariances is correct and is a real result: the prefactor identity works, the reciprocity of the noise is preserved, and the MSD comparisons give supporting evidence at finite ε. The most load-bearing assumption, however, is the replacement of the exact discrete noise distributions by Gaussian noise with matched first two moments. That assumption is not validated by MSD agreement alone, because MSD is a bulk dynamical observable that is insensitive to the rare, sparse events that control an absorbing-state transition. The paper's own text concedes that the stochastic approximation degrades for larger kick sizes, so the approximation is not uniformly valid; its safety in the small-ε limit is exactly what needs to be demonstrated for the critical φc claim. The fixed-exponent FSS analysis is a secondary weakness — it tests consistency with Manna exponents rather than measuring them — but the Gaussian substitution is more fundamental because it underlies the analytic equivalence and the mapping to the SDE. I do not think this warrants rejection: the replacement is plausible, the numerics are suggestive, and the reader's conditional verdict already reflects the need for independent verification. My concern therefore does not change the verdict; it sharpens the specific check that should be run first.","tokens_in":14171,"tokens_out":9533,"duration_ms":107351,"concrete_test":"Simulate the exact pairwise BRO and SGD algorithms alongside their Gaussian stochastic approximations, Eqs. 4 and 6, at matched parameters b_f = 3/4, α = 4ε/3, ε = 10^-3 (2R), for N = 10^5 particles on the same grid of φ near 0.64 and with the same initialization protocol. Perform free-exponent finite-size scaling fits — without fixing β, ν∥, or ν⊥ — on the steady-state activity and relaxation time for each of the four processes. If the fitted φc or exponents of either stochastic approximation differ from those of its discrete parent by more than the fit uncertainty, the Gaussian replacement changes the critical behavior and the Eq. 6 = Eq. 4 equivalence does not by itself establish the central claim. Repeat at ε = 5 × 10^-4 to check whether any discrepancy shrinks with ε.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. 2 replaces the discrete per-pair BRO kick, uniform on [0, ε], with a Gaussian of the same mean and covariance, and the same replacement is used for the Bernoulli pair-selection noise when deriving Eq. 6. This Gaussian approximation is precisely what makes Eq. 6 match Eq. 4 at α = ε/b_f, b_f = 3/4, so the claimed BRO–SGD equivalence inherits its content from this substitution. The paper validates the substitution by comparing mean-square displacements at small ε (Fig. 5a, Fig. S1), but MSD agreement at φ = 0.63 and k = 10^4 does not constrain the absorbing-state transition: near criticality activity is sparse, a particle typically receives very few kicks before freezing, and no central-limit averaging over many kicks can be invoked. The true noises have different higher cumulants (uniform vs. two-point Bernoulli, with the Bernoulli noise skewed), and if any higher cumulant is relevant at the Manna fixed point, the exact discrete algorithms and their Gaussian approximations need not share the same φc or the same universality class. The covariance calculation itself is correct; the load-bearing assumption is that the Gaussian replacement is harmless for critical behavior, and that assumption is currently untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies a minimal model of 'neural manifold packing' in which stochastic gradient descent (SGD) is applied to spherical particles with short-ranged linear repulsive interactions, and compares it with biased random organization (BRO), a nonequilibrium absorbing-state model. The authors derive Gaussian approximations to pairwise BRO and to their energy-based SGD by matching the mean and covariance of the discrete updates. They show that, at learning rate α=ε/b_f and batch fraction b_f=3/4, the two Gaussian-approximated processes coincide (Eqs. 4 and 6), and similarly for individual-particle BRO and random coordinate descent (RCD). Numerical simulations report mean-square displacement agreement at small kick sizes, critical packing fractions φc→0.64 as ε,α→0, finite-size scaling consistent with the Manna universality class for both SGD and RCD at various b_f (including b_f=1, i.e., deterministic gradient descent), and a flatness bias for SGD above φc.","tokens_in":14422,"tokens_out":11485,"duration_ms":113715,"significance":"The conceptual bridge between a machine-learning optimization algorithm and absorbing-state critical phenomena is attractive. The mean-covariance computations in the Appendix are transparent, and I verified the exact parameter matching at α=ε/b_f, b_f=3/4; within the Gaussian approximation, this is a genuine parameter-free derivation. The external benchmark against the known BRO value φc≈0.64 is a nice sanity check. However, the paper's central claims rest on the assumption that replacing the true discrete noises by Gaussians does not alter the absorbing-state critical behavior, and this assumption is currently tested only by MSD agreement away from criticality. The Manna-universality claim is also weakened by the a priori fixing of exponents in the finite-size scaling analysis. If the proposed near-critical tests are passed, the paper would make a valuable contribution to the statistical-mechanics understanding of SGD-like dynamics.","major_comments":[{"comment":"The central equivalence between BRO and SGD is derived only after replacing the true discrete noises by Gaussian noises with matched mean and covariance; Eq. (6) equals Eq. (4) exactly only for these Gaussian-approximated processes. The validation in Fig. 5(a) and Fig. S1 compares mean-square displacements at φ=0.63 for small kick sizes, which does not probe the absorbing-state transition. Near φc activity is sparse and each particle receives very few kicks before freezing, so no central-limit argument justifies the replacement; the higher cumulants of the uniform (BRO) and Bernoulli (SGD) noises, which differ, could in principle change φc or the universality class. A concrete test would be to compare the exact BRO/SGD dynamics with their Gaussian approximations for the steady-state activity, survival probability, and finite-size scaling near φc at the same small ε; without such a test the claimed equivalence is conditional on an unverified assumption.","section":"Stochastic approximations, Eqs. (2) and (6)"},{"comment":"The claim that SGD and RCD belong to the Manna universality class is supported by finite-size scaling analyses in which the Manna exponents β=0.84, ν∥=1.08, and ν⊥=0.59 are fixed a priori and only φc is fitted; the text states that this is done 'assuming the Manna exponents'. This is a consistency check, not a measurement, because imposing the exponents cannot distinguish Manna from nearby universality classes. A stronger test would fit β and ν∥ as free parameters, or at least compare the collapse quality against alternative exponent sets. The paper should either soften the claim to 'consistent with Manna exponents' or provide such a free-exponent analysis.","section":"Critical behavior, Fig. 3 and Manna universality"},{"comment":"For b_f=1.0, SGD and RCD reduce to deterministic gradient descent, which the paper itself notes is not an absorbing-state model because it 'absorbs' on both sides of φc. The finite-size scaling shown in Fig. 3(c,d) for b_f=1.0 is therefore not obviously well-defined: for φ>φc the steady state is a static minimum, not an active fluctuating steady state, and the paper does not specify how fa∞ and τr are computed in that case. The claim that the zero-noise limit is also in the Manna universality class needs either a precise definition of the observables for GD or a separate argument that the deterministic limit is a well-defined absorbing-state dynamics.","section":"Critical behavior, b_f=1.0 paragraph"},{"comment":"The theoretical justification of the flatness measure, ΔV ∝ Tr(H), assumes a second-order smooth potential, but the model potential U(r)=1/2(2R−r) for r<2R is piecewise linear: its Hessian is zero in the interior of the overlap region and singular at the contact point r=2R. The relationship ΔV ∝ Tr(H) therefore does not follow for this model, and contributions from the potential's kink may dominate the energy fluctuation. The authors should either use a smooth potential for the flatness comparison or justify the flatness measure directly without the Hessian expansion. In addition, Fig. 4(b) shows no error bars for the reported energy fluctuations.","section":"Flatness analysis, Fig. 4 and Eq. (S12)"}],"minor_comments":[{"comment":"The sentence 'the stochastic approximations of individual-particle BRO and SGD are equivalent' should read 'individual-particle BRO and RCD', since Eq. (A8) is matched to the individual-particle BRO approximation, while SGD is matched to pairwise BRO in Eq. (6).","section":"Appendix A, final paragraph"},{"comment":"The active-particle condition contains '|xj_k−xj_k|<2R'; this should be '|xj_k−xi_k|<2R'.","section":"Algorithm 2, Supplementary Material"},{"comment":"Equation (S5) is missing a minus sign on the drift term; it should read dx(t) = −ε/τ ∇iV dt + ..., consistent with Eq. (5) in the main text.","section":"Supplementary Material, Eq. (S5)"},{"comment":"The text defines b_f as the 'fraction of pairs in a batch', but Algorithm 3 selects each active pair independently with probability b_f, so the realized batch fraction is random; please clarify whether b_f denotes the expected fraction or a fixed batch size.","section":"Methods, definition of b_f"},{"comment":"The convergence of φc to φRCP as α→0 is presented without error bars or a specified fitting procedure for the extrapolation; reporting the fitted values and their statistical uncertainties would strengthen the claim.","section":"Fig. 5(b)"},{"comment":"The paper would benefit from a data/code availability statement, given that all central results are numerical.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript is a revised arXiv posting; I see no issues with attribution, and the citation practice to the BRO and SGD literature is appropriate. The connection to neural manifolds is speculative, but the paper is framed as a minimal model, which is acceptable for a statistical mechanics journal. My recommendation of major revision is driven by a single methodological gap—validating the Gaussian approximation near criticality—which is fixable with additional simulations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe one thing to know: this paper offers the most direct mapping I've seen between SGD and an absorbing-state model, and the numerical evidence that SGD sits in the Manna universality class is reasonably convincing. The caveat is that the analytic equivalence goes through a Gaussian substitution that is validated only in a regime that does not test critical behavior.\n\nWhat's new: the explicit match between pairwise BRO and SGD at batch fraction b_f=3/4 and learning rate α=ε/b_f. Eq. 6 has the same form as Eq. 4 after that substitution. The derivation by mean-covariance matching is clean and the algebra checks out. The claim that RCD maps to individual-particle BRO is likewise new. The FSS on the actual SGD and RCD algorithms across batch fractions, including the b_f=1 (GD) limit, showing Manna scaling and φc→0.64, is a solid piece of numerics. The flat-minima result—SGD favors flat minima while RCD favors sharp minima—is a nice qualitative corollary.\n\nWhere the paper is soft: the Gaussian replacement of the discrete noise (uniform kicks in BRO, Bernoulli pair selection in SGD) is the load-bearing step in the proof that the two processes are equivalent. The paper tests it by comparing MSDs at φ=0.63 and k=10^4, which is below criticality and says little about whether the absorbing-state transition is preserved. Near criticality activity is sparse and each particle gets few kicks, so the central-limit intuition doesn't apply. If higher cumulants of the noise are relevant at the Manna fixed point, the Gaussian equivalence could miss the true critical behavior. That said, the numerical FSS is performed on the original algorithms, not the Gaussian approximation, so the main numerical conclusion—that SGD is in the Manna class—doesn't ride on the approximation alone. Still, the analytic correspondence would be stronger with some argument that the noise cumulants beyond covariance are irrelevant, or with a direct comparison of the critical behavior of the original and approximated processes.\n\nTwo more minor gripes: the FSS assumes the Manna exponents and fits only φc; a free-exponent fit would be more persuasive. And there are no error bars on the key figures, nor code/data artifacts, which makes it harder to assess the collapse quality.\n\nWho this is for: stat mech people working on absorbing transitions and anyone interested in the noise-dependence of SGD. The neural manifold analogy is a wrapper; the core is a clean minimal model.\n\nRecommendation: I'd send it to review. The central claim is interesting and mostly supported, and the Gaussian-approximation concern is addressable with additional numerics or a reanalysis. A serious referee would catch issues that the authors can fix.","headline":"A clean new mapping between SGD and BRO, with solid numerics, but the Gaussian substitution is the weak point in the analytic equivalence.","tokens_in":14957,"tokens_out":3568,"would_cite":true,"duration_ms":36375,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Stochastic gradient descent reduces to biased random organization near a critical packing fraction of about 0.64.","keywords":["stochastic gradient descent","biased random organization","absorbing phase transition","Manna universality class","random close packing","neural manifold packing","multiplicative noise"],"falsifier":"Compare the critical packing fraction and the steady-state activity scaling of SGD and BRO at a fixed small $\\alpha=\\epsilon/b_f$, $b_f=3/4$, but replace the Bernoulli batch selection with a heavy-tailed or skewed selection rule that keeps the same mean and covariance; if $\\phi_c$ or the measured exponents move away from $0.64$ and the Manna values, the higher noise cumulants are relevant and the claimed equivalence would fail.","tokens_in":13939,"feed_emoji":"📦","tokens_out":9892,"duration_ms":85098,"temperature":0.7,"pith_summary":"The paper claims that stochastic gradient descent—the workhorse optimizer of deep learning—behaves, in a minimal model of sphere packing, exactly like biased random organization (BRO), a nonequilibrium absorbing-state model. The two processes have unrelated sources of randomness: BRO's kicks are random in size, while SGD's randomness comes from randomly sampled batches of pairwise interactions. Yet the paper shows that under a Gaussian (mean-covariance) approximation of their update noise, the two stochastic processes coincide when the learning rate and batch fraction are set to $\\alpha=\\varepsilon/b_f$ and $b_f=3/4$. As a consequence, SGD, BRO, and random coordinate descent all converge to the same critical packing fraction $\\phi_c\\approx 0.64$ as the kick size or learning rate goes to zero, and all belong to the Manna universality class near criticality. Above the critical point, SGD's batch noise biases it toward flat minima of the energy landscape, connecting the physics model to an experimentally observed property of neural-network training.","feed_headline":"SGD matches an absorbing-state model at packing fraction 0.64","feed_subtitle":"A mean-covariance match shows learning-rate noise and batch selection noise play the same role near criticality.","key_machinery":"The central object is the mean-covariance (Gaussian) approximation of the discrete update noise: each random kick (BRO) or random batch selection (SGD) is replaced by a Gaussian with matching first and second moments, turning both processes into stochastic differential equations with anisotropic multiplicative noise. For pairwise BRO this is $dx_i = -\\frac{\\epsilon}{\\tau}\\nabla_i V\\,dt + \\frac{\\epsilon}{\\sqrt{3\\tau}}\\sum_j\\sqrt{\\Lambda_{ji}}\\,dW_{ji}$, and for SGD it is the analogous equation with prefactors $\\alpha b_f$ for the drift and $\\alpha\\sqrt{b_f-b_f^2}$ for the noise. The two SDEs match exactly under the parameter map $\\alpha=\\epsilon/b_f$, $b_f=3/4$, so the approximation is the load-bearing identity that makes the dynamical equivalence a theorem about the two original algorithms.","core_discovery":"Starting from the pairwise BRO update, in which each overlapping pair is displaced by equal-magnitude random kicks along their line of centers, the authors decompose each kick into its mean and its fluctuation and replace the latter by Gaussian noise with the same covariance. The mean is the negative gradient of a linear repulsive potential $U(r)=\\frac12(2R-|r|)$, and the fluctuation becomes an anisotropic multiplicative noise term $\\frac{\\epsilon}{\\sqrt{3}}\\sum_j \\sqrt{\\Lambda_{ji}}\\,\\xi_{ji}$. Applying the same mean-covariance decomposition to an energy-based SGD that selects active pairs with probability $b_f$ and moves them by $\\alpha$-sized gradient steps yields exactly the same stochastic approximation when $\\alpha=\\epsilon/b_f$ and $b_f=3/4$. The paper therefore establishes that the two dynamics are equivalent in distribution at small kick sizes, and verifies numerically that their mean square displacements, structure factors, and finite-size scaling agree. At the critical point both schemes reach the random-close-packing fraction $\\phi_c\\approx0.64$ as $\\epsilon,\\alpha\\to0$, and the activity and relaxation-time exponents match the Manna universality class values ($\\beta=0.84$, $\\nu_\\parallel=1.08$) regardless of batch size.","pith_inferences":["Because only the mean and covariance of the noise are matched, the equivalence suggests that higher-order noise statistics (skewness, heavy tails) are irrelevant for the absorbing transition; the paper does not prove this, and it is a testable prediction of the stochastic-approximation framework.","The explicit SDE approximations allow the authors' conjecture of mean-field behavior for $d\\ge4$ to be tested in simulation without running the full discrete algorithms, since integrating the SDEs at $d=4$ and $d=5$ would show whether the Manna exponents cross over.","In practical representation learning, the result implies that batch size and learning rate are physical control parameters: they place the learning dynamics on one side or the other of an absorbing transition between fully separated and partially overlapping neural manifolds, a perspective the paper opens but does not develop into algorithmic advice."],"forward_implications":["As the learning rate goes to zero, SGD's critical packing fraction rises toward the random-close-packing value $\\phi_c \\approx 0.64$, regardless of batch fraction.","Near the transition, the steady-state activity and relaxation time of SGD and RCD collapse onto Manna-universality scaling curves, including at $b_f=1$ where SGD reduces to gradient descent.","Because the stochastic approximation matches exactly, the equivalence implies that the noise statistics of minibatch selection and BRO's kick-size statistics play identical roles in the critical region.","Above the critical point, SGD with small batch sizes settles into flatter minima than gradient descent, whereas RCD settles into sharper ones."],"supporting_citations":[{"why":"Reference data for BRO's critical point at the random-close-packing fraction $\\phi_c\\approx 0.64$ and its Manna-class critical behavior, which SGD is claimed to reproduce.","marker":"[37]"},{"why":"Defines the pairwise and individual-particle BRO update rules that the paper approximates and compares against.","marker":"[39]"},{"why":"Supplies the form of the pairwise BRO kick update (with reciprocal uniform noise) used in the paper's Eq. 1.","marker":"[49]"},{"why":"Provides the stochastic-approximation technique of replacing discrete updates by SDEs with Gaussian noise, which the paper applies to both BRO and SGD.","marker":"[17]"},{"why":"Models SGD as a continuous-time SDE with Gaussian noise, the same representation underlying the equivalence.","marker":"[51]"},{"why":"One of the references for the Manna-class values $\\beta=0.84$ and $\\nu_\\parallel=1.08$ that the finite-size scaling analysis assumes and confirms.","marker":"[55]"},{"why":"Additional reference for the Manna critical exponents used in the scaling collapse.","marker":"[56]"},{"why":"The Python package used to perform the finite-size scaling fits that yield the reported $\\phi_c$ values and universality-class verification.","marker":"[58]"}],"fun_headline_variants":["SGD matches biased random organization at critical packing 0.64","Stochastic gradient descent exhibits absorbing phase transition","SGD dynamics equivalent to BRO for small learning rates","SGD and BRO share Manna universality near 0.64 packing","Neural manifold packing: SGD and BRO share criticality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Gaussian mean-covariance approximation of the discrete update noise is assumed to preserve the absorbing-state critical behavior, even though only the first two noise moments are matched and the paper validates the approximation numerically at small kick sizes rather than proving that higher cumulants are irrelevant.","fun_headline_variants_meta":{"raw":{"variants":["SGD matches biased random organization at critical packing 0.64","Stochastic gradient descent exhibits absorbing phase transition","SGD dynamics equivalent to BRO for small learning rates","SGD and BRO share Manna universality near 0.64 packing","Neural manifold packing: SGD and BRO share criticality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000648,"raw_usage":{"total_tokens":3033,"prompt_tokens":1064,"completion_tokens":1969,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":1886}},"tokens_in":680,"tokens_out":1969,"duration_ms":14274,"temperature":1.0,"reasoning_tokens":1886,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:06:16.040629+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the critical packing fraction and the steady-state activity scaling of SGD and BRO at a fixed small $\\alpha=\\epsilon/b_f$, $b_f=3/4$, but replace the Bernoulli batch selection with a heavy-tailed or skewed selection rule that keeps the same mean and covariance; if $\\phi_c$ or the measured exponents move away from $0.64$ and the Manna values, the higher noise cumulants are relevant and the claimed equivalence would fail.","supporting_citations":[{"cited_title":"Wilken, R","cited_arxiv_id":null,"evidence_quote":"Defines the pairwise and individual-particle BRO update rules that the paper approximates and compares against."},{"cited_title":"Ness and M","cited_arxiv_id":null,"evidence_quote":"Supplies the form of the pairwise BRO kick update (with reciprocal uniform noise) used in the paper's Eq. 1."},{"cited_title":"Mandt, M","cited_arxiv_id":null,"evidence_quote":"Models SGD as a continuous-time SDE with Gaussian noise, the same representation underlying the equivalence."},{"cited_title":"Martiniani, P","cited_arxiv_id":null,"evidence_quote":"One of the references for the Manna-class values $\\beta=0.84$ and $\\nu_\\parallel=1.08$ that the finite-size scaling analysis assumes and confirms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Additional reference for the Manna critical exponents used in the scaling collapse."},{"cited_title":"Sorge, Zenodo: A scientific Python package for finite- size scaling analysis (2015)","cited_arxiv_id":null,"evidence_quote":"The Python package used to perform the finite-size scaling fits that yield the reported $\\phi_c$ values and universality-class verification."}],"review_version":1}