Pith. sign in

REVIEW 3 major objections 5 minor 22 references

Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

T0 review · 3 major / 5 minor · reviewed 2026-07-31 · deepseek-v4-flash

Pith's one-line read Grokking is a slow dissipative relaxation on a weight-decay clock; the paper derives its exact rate and isolates the subspace that drives it.

desk verdict Strong exact-solvable core for grokking's slow tail; the nonlinear benchmark is the main weak spot, but the paper is honest about it and deserves serious refereeing. read the letter →

arxiv 2607.23967 v1 pith:JJPFHHSR submitted 2026-07-27 cs.AI cs.LG

classification cs.AIcs.LG
keywords grokkingdelayedgeneralizationweightdecayheavy-ballmomentumempiricalnullspacepopulationrisksoftNoetherlawneuraltangentkernel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to prove that grokking — the delayed jump in generalization long after training loss converges — has a specific, computable late-time cause in models trained with heavy-ball momentum and weight decay. It identifies a 'grokking subspace': parameter directions that are invisible on the training set but visible at the population level. Because training loss cannot see these directions, weight decay is the only force pulling them back, and that soft breaking of a symmetry sets a slow exponential clock whose rate is exactly ρ₀, with grokking time ≈ (1−β)/(ηλ) log(‖Π_g ξ₀‖/ε). The paper shows only this subspace contributes to the slow asymptotic decay of population risk, while the training loss is provably blind to it, and verifies all identities without fitted parameters in a solvable synthetic model and in a modular-addition benchmark. If right, it turns grokking from a mysterious phase transition into a rate phenomenon on a known clock, with testable interventions.

What carries the argument

The grokking subspace S = Ker G_N ∩ (Ker G_pop)^⊥ — the metric-chosen representative of the quotient of empirical-null directions by population-null directions — is the central object. The carrying identity is the soft Noether law for translation charges: dQ_b/dt = −(γ/m)Q_b − λ q_b, which is exactly the damped-oscillator equation of the slow null coordinate. Its exact discrete counterpart has characteristic root ρ₀ = (a₀ + √(a₀² − 4β))/2, a₀ = 1 + β − ηλ, yielding the iteration clock k_grok ≈ (1−β)/(ηλ) log(‖Π_g ξ₀‖/ε). The paper derives a full rate hierarchy — slow ρ₀^k, transverse √β^k, angular β^k — and proves the training loss is blind to the slow root.

What would settle it

A concrete check: in the synthetic model, violate Assumption 19 by raising λ toward the stability boundary or lowering a transverse eigenvalue until ϱ ≥ ρ₀², then measure the population-risk tail; if it still decays as Bρ₀^{2k} with no faster contamination, the separation assumption is not load-bearing, whereas any visible ϱ^k contamination confirms the theorem's stated scope.

Watch

Extended reading notes

Core claim

The paper's central claim, stated on its own terms, is that grokking in the exactly solvable setting — a model linear in its parameters, squared loss, full-batch heavy ball with coupled L₂ regularization — is a slow dissipative relaxation along a distinguished subspace, not a representational phase transition. The empirical Gram matrix and its population counterpart have nested null spaces, N_pop ⊆ N_N, and the quotient's orthogonal representative S = N_N ∩ (N_pop)^⊥ is the grokking subspace: directions along which training predictions are frozen to first order while population-level predictions still move. Translations along S are exact symmetries of the training loss; weight decay softly b

Load-bearing premise

The load-bearing premise is tail separation: every faster mode (the fast null root and all transverse modes) must decay strictly faster than the square of the slow root, ϱ < ρ₀², since otherwise those modes contaminate the ρ₀^{2k} tail and the claim that only the grokking subspace drives the slow asymptotic decay is no longer guaranteed.

Editorial extensions

If this is right

  • Grokking time in the solvable regime collapses onto the exact prediction log(‖q₀‖/ε)/(−log ρ₀) across (λ, η, β), with the weak-regularization law (1−β)/(ηλ) as its leading approximation.
  • Only the grokking subspace contributes to the slow asymptotic decay of the population risk; population-null directions decay at the same parameter rate but are invisible in function space, and transverse directions decay strictly faster.
  • The population-risk tail is ρ₀^{2k} when the compatibility condition Π_g ∇L_pop(θ_eq) = 0 holds and ρ₀^k when it fails; since the linear coefficient A can be negative, the population risk can approach its limit from below.
  • Rescaling the grokking component at a post-interpolation time leaves every training prediction unchanged and shifts the grokking iteration count by log c/(−log ρ₀) — a causal, parameter-free intervention prediction.
  • Coupled L₂ regularization accelerates grokking by factor (1−β)⁻¹ relative to decoupled weight decay at the same (η, λ, β), a direct optimizer-dependent prediction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not claim the loss tail explains the abrupt accuracy jump; a separate margin argument is left open. A natural follow-up is to convert the predicted population-risk tail into a threshold time for accuracy, which would make the theory directly comparable to the curves that define grokking.
  • Because the rate ρ₀ is independent of dataset size while subspace dimensions, initial amplitudes, and the compatibility condition depend on it, the theory predicts dataset size shifts the log factor and the A/B coefficients but not the exponential clock — a separation that could be tested by N-sweeps.
  • The coupled-versus-decoupled weight decay factor (1−β) implies momentum changes the semantics of weight decay, not just the speed of training. Optimizer ablations (SGD+L₂ vs heavy-ball+L₂ vs SGDW) are therefore a sharper test of the mechanism than any single grokking curve.
  • Stochastic gradients project noise onto the null space and could compete with the λ-drift; a stochastic version of the soft Noether law, which the paper lists as a priority, would show whether minibatch training renormalizes the clock or preserves its form.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies grokking in linear-in-parameter models trained with full-batch heavy-ball momentum and coupled L2 weight decay. It derives an exact modal decomposition of the training recurrence, identifies the empirical null space as carrying prediction-preserving translation and rotation symmetries, and defines the grokking subspace S = Ker G_N ∩ (Ker G_pop)^⊥ as the population-active part of the empirical null space. Under explicit assumptions including weak regularization and a tail-separation condition (Assumption 19), Theorem 20 shows that the signed excess population risk relative to the regularized equilibrium decays as A ρ0^k + B ρ0^{2k} + R_k with |R_k| ≤ C ϱ^k, where only the S-component enters the slow tail; the training loss is provably blind to the slow rate ρ0. The paper predicts a grokking time kgrok ≈ ((1−β)/(ηλ)) log(||Πg ξ0||/ε), distinguishes coupled L2 from decoupled weight decay, provides a local nonlinear extension with explicit error floor, and reports parameter-free verification in a synthetic model plus a modular-addition benchmark with measured log–log scaling slope 1.01.

Significance. If the results are taken as stated, this is a valuable exactly solvable reference mechanism for post-interpolation delayed generalization. The exact linear theory is rigorously derived and unusually complete: the mode decomposition and tail calculation are checkable, the synthetic verification uses no fitted parameters, and the nonlinear analysis displays its remainder terms and error floor rather than hiding them. The distinction between coupled and decoupled weight decay and the intervention predictions (P2–P4) are genuinely falsifiable. The main limitation is that the exact claims are confined to a narrow regime, and the bridge to the nonlinear benchmark is more qualitative than the abstract suggests.

major comments (3)
  1. [§5, Assumption 19 and Eq. (34)] Theorem 20's central claim that 'only the grokking subspace contributes' to the slow asymptotic population-risk tail is conditional on the tail-separation Assumption 19: ϱ := max(ρ1, max_{j:h_j>0}|ρ±_j|) < ρ0^2. The proof in Appendix B.3 makes clear that R_k is only O(ϱ^k), so if this assumption fails, transverse or fast-null modes can dominate the ρ0^{2k} term, and the 'only S' statement no longer follows. Assumption 19 is not implied by stability (Assumption 18), and it is not verified in the modular-addition benchmark. In fact, Section 7.5 reports that the population-visible empirical-null component measured with projectors frozen at k=2000 decays several times more slowly than ρ0, indicating that the spectral separation can fail in a relevant nonlinear setting. Please either verify ϱ<ρ0^2 in the benchmark using the empirical Jacobian spectrum, or qualify the abstract and Theorem 20 a
  2. [§7.5, Figure 5(c)] The nonlinear benchmark validates (P1) by fitting a log–log slope of 1.01 for the iteration at test accuracy 0.9 versus (1−β)/(ηλ). But the theory's kgrok is defined as a parameter-space or population-risk-excess threshold (Corollaries 5 and 22), and the paper explicitly leaves open the margin argument that converts a loss tail into the abrupt accuracy jump (Section 9). Since the quantity plotted in Figure 5(c) is an accuracy threshold, the benchmark does not test the predicted prefactor or the exact ρ0 rate; it tests only the exponent. Please report a loss-threshold grokking time, or otherwise justify that the accuracy-threshold time inherits the same implied constant.
  3. [§7.5, Figure 5(b) and Abstract] The abstract says that in modular addition 'the late-time relaxation agrees closely with the theoretical clock.' The measured test-loss excess tail decays at 4.5×10^-4 per step against the parameter-free prediction −log ρ0 = 6.0×10^-4, a 25% discrepancy, while the population-visible empirical-null component with frozen projectors decays several times more slowly. The paper attributes this to Jacobian rotation, which is a plausible and honest explanation, but it means the benchmark does not quantitatively confirm the weight-decay clock; it only shows an exponential tail of the same order. The claim should be tempered, or the benchmark should include a check of the frozen-projector assumptions over the fitted window.
minor comments (5)
  1. [Abstract and §5] The abstract's 'only this subspace contributes' should be accompanied by a pointer to the exact-regime and spectral-separation assumptions, since the statement is conditional in the theorem.
  2. [§2.3, Eq. (7)] The definition of S as an orthogonal representative of the quotient depends on the Euclidean metric; the paper acknowledges this, but Figure 1 and several prose passages could more consistently say 'the chosen metric representative.'
  3. [§7.4, E3] The intervention experiment rescales both θ and the momentum buffer to stay on the slow branch. This is correct, but it would help to state explicitly that rescaling θ alone excites the fast null root ρ1 and therefore the measured slope would not be the predicted one; the text mentions this only in passing.
  4. [Appendix B.3] In the proof of Theorem 20, the handling of double roots ('absorbed by enlarging ρ̄ infinitesimally within Assumption 19') is correct but should be written as a limiting argument to avoid the impression that ρ̄ itself can be changed arbitrarily.
  5. [§7.5] The modular-addition results are single-seed and single-architecture; the paper says this in the limitations, and the proposed protocol in Section 7.6 is appropriate. It would strengthen the paper to include at least one of the planned multi-seed checks for the rate measurement.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central rates and tail coefficients are derived in closed form from the exact heavy-ball recurrence, and the fitted benchmark slope is a test statistic rather than a fitted input.

full rationale

The derivation chain is self-contained. Theorem 4 solves the linear heavy-ball recurrence exactly: the slow root ρ0 is a function of η, λ, β via Equation (16), and Corollary 5's grokking count follows from that root and the initial null-block component; no term in Eq. (20) is fitted to the quantity it predicts. Theorem 20 likewise expands the exactly quadratic population loss in the exact eigenmodes; the coefficient B = ½ s0^T Gpop s0 and the compatibility condition Eq. (37) are algebraic consequences of the definitions of Gpop, NN, Npop and S, not data-dependent fits. The synthetic experiments E1–E3 are parameter-free identity checks of these closed forms, and the modular-addition benchmark tests the scaling law by measuring a log–log slope (1.01) against the predicted exponent 1; that fitted slope is a diagnostic, not a parameter fed back into the predicted rate, which is computed from (η,λ,β). Assumption 19 is a genuine separation condition: the paper states it as an assumption, notes it is not implied by stability, and acknowledges that the frozen-projector description degrades in the modular benchmark (Section 7.5). A conditional theorem with an explicit, possibly-failing hypothesis is a robustness limitation, not circularity. There are no load-bearing self-citations: the relevant background (Boursier et al. 2025, Xu et al. 2026, etc.) is used only for comparison, and the paper's uniqueness claims are its own proofs. No step in the derivation reduces by construction to its own conclusion.

Assumptions & free parameters 0 free parameters · 9 assumptions · 1 invented entities

The central claim rests on the linear model and explicit Assumptions 16–19; no parameters are fitted to data—all rates and coefficients are computed from η, λ, β and the geometry. The nonlinear extension adds local linearity assumptions with error floors. The 'grokking subspace' is an invented mathematical object with falsifiable predictions.

assumptions (9)
  • domain assumption Linear model f(x;θ)=φ(x)^T θ (Section 3, Eq. 9)
    The exact theory applies only to models linear in parameters; this is the core solvable regime.
  • domain assumption Squared loss, giving exactly quadratic empirical and population risks (Section 2.1, Eq. 1)
    Required for the exact mode decomposition and quadratic tail calculation.
  • domain assumption Full-batch heavy-ball update with coupled L2 regularization (Section 2.2, Eq. 3)
    The optimizer recurrence is the object of analysis; decoupled weight decay yields a different clock (Proposition 6).
  • ad hoc to paper Assumption 18: ηλ < (1−√β)^2 and η(hj+λ) ≤ (1+√β)^2 for all j
    Ensures the strict rate hierarchy ρ0 > |ρ±_j| and excludes overdriven oscillatory modes near the stability boundary.
  • ad hoc to paper Assumption 19: ϱ := max(ρ1, max_{j:hj>0} |ρ±_j|) < ρ0^2
    Tail separation needed for the clean statement that only the grokking component contributes to the ρ0^{2k} tail.
  • domain assumption Assumption 17: Npop ⊆ NN (holds a.s. for i.i.d. samples by Lemma 3)
    Ensures the orthogonal splitting RP = N⊥_N ⊕ S ⊕ Npop.
  • domain assumption Assumptions 25–28: local linearization, constant rank, population-side local linearity, and spectral gap/Jacobian regularity
    Basis of the local nonlinear extension in Section 6; yields the error floor in Proposition 29.
  • ad hoc to paper Compatibility condition Πg∇Lpop(θeq)=0 (Eq. 37, Remark 21)
    Required for the clean ρ0^{2k} law in Theorem 20(ii); holds for isotropic population geometry or weak-regularization-suppressed linear term.
  • ad hoc to paper Realizable teacher θT ∈ N⊥_N for Corollary 23 (Remark 21a)
    Needed to connect relaxation toward θeq with monotone approach to the population optimum.
invented entities (1)
  • Grokking subspace S = Ker G_N ∩ (Ker G_pop)^⊥ independent evidence
    purpose: Identifies parameter directions that are invisible on training predictions yet visible at the population level; claimed to be the sole carrier of the slow population-risk tail.
    The subspace is measurable from data (Gram null spaces) and its dynamics are verified parameter-free in the synthetic model; the gauge-control experiment (Figure 4b) gives an operational train/population signature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries." pith.science (2026). https://pith.science/paper/JJPFHHSR

@misc{pith2026260723967,
  author       = {Pith},
  title        = {Pith review of: Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JJPFHHSR}},
  note         = {Machine review of arXiv:2607.23967}
}
abstract

Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight decay, together with a locally quadratic extension to nonlinear neural networks. Our analysis reveals a distinguished population-active component of the empirical null space, which we call the grokking subspace. Along this subspace, the training predictions remain unchanged, leaving weight decay as the sole restoring force and giving rise to a slow dissipative relaxation governed by an exact discrete-time and continuous-time law. We show that only this subspace contributes to the slow asymptotic decay of the population risk and derive explicit iteration-scale predictions for the grokking time, recovering the familiar $(1-\beta)/(\eta\lambda)$ scaling in the weak-regularization regime. The theory further predicts distinct effects of optimizer choice, distinguishing coupled $L_2$ regularization from decoupled weight decay, and yields causal predictions for interventions that modify the grokking component. We verify all theoretical identities without fitted parameters in a synthetic model where every subspace and relaxation rate is computable in closed form. We further observe genuine delayed generalization in modular addition, where the measured delay follows the predicted scaling and the late-time relaxation agrees closely with the theoretical clock.

Figures

Figures reproduced from arXiv: 2607.23967 by the authors.

Figure 1
Figure 1. Kernel decomposition and the three clocks. Left: the parameter space splits into [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. E1 on the synthetic model (P = 60, d = 20, N = 8, η = 0.05, λ = 2 × 10−3 , β = 0.95). (a) Component norms and rotation charge against iteration; dotted lines are the parameter-free predictions β k/2 , ρ k 0 , and β k . (b) Training loss and excess population risk for the compatible and generic teachers; dotted curves are the fully parameter-free predictions Bρ2k 0 and |Aρk 0 + Bρ2k 0 | with A, B computed from Equati… view at source ↗
Figure 3
Figure 3. E2 scaling sweep over 36 configurations of ( [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: E3 interventions. (a) Rescaling the grokking component at [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]
Figure 5
Figure 5. Figure 5: Modular-addition benchmark (p = 17, width 128, quadratic activation, squared loss, heavy ball with coupled L2; single seed). (a) Genuine delayed generalization at η = 0.2, λ = 3 × 10−4 , β = 0.9. (b) Test-loss excess relative to its late plateau, the population-visible…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 2 linked inside Pith

  1. [1]

    2022 , eprint =

    Alethea Power and Yuri Burda and Harri Edwards and Igor Babuschkin and Vedant Misra , title =. 2022 , eprint =

  2. [2]

    Nolte and Eric Michaud and Max Tegmark and Mike Williams , title =

    Ziming Liu and Ouail Kitouni and Niklas S. Nolte and Eric Michaud and Max Tegmark and Mike Williams , title =. 2022 , booktitle =

  3. [3]

    Michaud and Max Tegmark , title =

    Ziming Liu and Eric J. Michaud and Max Tegmark , title =. The Eleventh International Conference on Learning Representations , year =. 2210.01117 , archivePrefix =

  4. [4]

    2023 , booktitle =

    Neel Nanda and Lawrence Chan and Tom Lieberum and Jess Smith and Jacob Steinhardt , title =. 2023 , booktitle =

  5. [5]

    2023 , eprint =

    Andrey Gromov , title =. 2023 , eprint =

  6. [6]

    Sutherland , title =

    Mohamad Amin Mohamadi and Zhiyuan Li and Lei Wu and Danica J. Sutherland , title =. Proceedings of the 41st International Conference on Machine Learning , volume =. 2024 , pages =

  7. [7]

    Baraniuk , title =

    Ahmed Imtiaz Humayun and Randall Balestriero and Richard G. Baraniuk , title =. Proceedings of the 41st International Conference on Machine Learning , volume =. 2024 , pages =

  8. [8]

    Proceedings of the 42nd International Conference on Machine Learning , volume =

    Alon Beck and Noam Itzhak Levi and Yohai Bar-Sinai , title =. Proceedings of the 42nd International Conference on Machine Learning , volume =. 2025 , pages =

Show all 22 references
  1. [9]

    Proceedings of the 42nd International Conference on Machine Learning , volume =

    Tikeng Notsawo Pascal Junior and Guillaume Dumas and Guillaume Rabusseau , title =. Proceedings of the 42nd International Conference on Machine Learning , volume =. 2025 , pages =

  2. [10]

    Saxe and James L

    Andrew M. Saxe and James L. McClelland and Surya Ganguli , title =. International Conference on Learning Representations , year =. 1312.6120 , archivePrefix =

  3. [11]

    Du and Wei Hu and Jason D

    Simon S. Du and Wei Hu and Jason D. Lee , title =. 2018 , booktitle =

  4. [12]

    Neural Computation , year =

    Amari, Shun-ichi , title =. Neural Computation , year =

  5. [13]

    Neural Tangent Kernel: Convergence and Generalization in Neural Networks , year =

    Arthur Jacot and Franck Gabriel and Cl\'. Neural Tangent Kernel: Convergence and Generalization in Neural Networks , year =. Advances in Neural Information Processing Systems , volume =

  6. [14]

    2025 , booktitle =

    Etienne Boursier and Scott Pesme and Radu-Alexandru Dragomir , title =. 2025 , booktitle =

  7. [15]

    2021 , booktitle =

    Hidenori Tanaka and Daniel Kunin , title =. 2021 , booktitle =

  8. [16]

    Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows , year =

    Sibylle Marcotte and R. Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows , year =. Advances in Neural Information Processing Systems , volume =

  9. [17]

    Keep the Momentum: Conservation Laws Beyond

    Sibylle Marcotte and R. Keep the Momentum: Conservation Laws Beyond. Proceedings of the 41st International Conference on Machine Learning , volume =. 2024 , pages =

  10. [18]

    Lee and Wei Hu , title =

    Kaifeng Lyu and Jikai Jin and Zhiyuan Li and Simon Shaolei Du and Jason D. Lee and Wei Hu , title =. 2024 , booktitle =

  11. [19]

    Gershman and Cengiz Pehlevan , title =

    Tanishq Kumar and Blake Bordelon and Samuel J. Gershman and Cengiz Pehlevan , title =. 2024 , booktitle =

  12. [20]

    Benjamin and David Klindt , title =

    Xingyu Zheng and Kyle Daruwalla and Ari S. Benjamin and David Klindt , title =. Proceedings of UniReps: the Second Edition of the Workshop on Unifying Representations in Neural Models , volume =. 2024 , pages =

  13. [21]

    2026 , eprint =

    Mingyue Xu and Gal Vardi and Itay Safran , title =. 2026 , eprint =

  14. [22]

    2026 , eprint =

    Truong Xuan Khanh and Truong Quynh Hoa and Luu Duc Trung and Phan Thanh Duc , title =. 2026 , eprint =

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.