REVIEW 2 major objections 2 minor 2 cited by
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent
T0 review · 2 major / 2 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read Fluctuations in the relative speeds of representation learning and readout calibration explain grokking and epoch-wise double descent.
desk verdict The representation-readout split organizes grokking and double descent under relative speeds but the metrics may not isolate the processes cleanly enough to support the attribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The representation-readout decomposition, which tracks representation quality with representational geometry and neural tangent kernels while tracking readout quality with linear probing.
What would settle it
Track representation quality and readout accuracy on held-out data throughout training on a grokking task; the central claim would be falsified if the two quantities do not show independent speed changes that align with the observed delay in test performance.
Extended reading notes
Core claim
Both representation learning and readout calibration are active throughout training, with the fluctuations of their relative speed giving rise to seemingly anomalous generalization dynamics. Applying the representation-readout decomposition to grokking across a wide range of tasks and architectures, we find that the readout is train-biased before grokking onset, and representation learning is gradual but not absent, contrary to the lazy-to-rich account. The framework further provides diagnostic signatures distinguishing spurious from genuine generalization: in a previously reported MNIST grokking example and an epoch-wise double descent example, apparent delayed or non-monotone generalizatio
Load-bearing premise
That representational geometry, neural tangent kernels, and linear probing cleanly separate representation learning from readout calibration without the measurements themselves being altered by the ongoing training dynamics.
Editorial extensions
If this is right
- The readout is train-biased before grokking onset.
- Representation learning is gradual but not absent during the period before grokking.
- Non-standard training recipes can induce representation degradation or readout misalignment that produces spurious delayed generalization.
- Genuine generalization is diagnosed when representation and readout metrics improve together rather than diverge.
Reading between the lines
- Interventions that selectively speed up representation learning while slowing readout calibration could shorten the grokking delay.
- The same decomposition may clarify other training phenomena in which train and test curves diverge, such as certain forms of overfitting or forgetting.
- Simple synthetic models could be constructed in which the two speeds are controlled independently to reproduce grokking without large networks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript claims that grokking and epoch-wise double descent arise from fluctuations in the relative speeds of two processes: representation learning in the encoder and readout calibration in the final classifier. Using representational geometry, neural tangent kernels, and linear probing across a range of tasks and architectures, it shows both processes remain active throughout training; the readout is train-biased before grokking onset while representation learning proceeds gradually (contrary to lazy-to-rich accounts). The framework also supplies diagnostic signatures that attribute some previously reported instances of these phenomena to representation degradation and readout misalignment induced by non-standard training recipes.
Significance. If the decomposition is shown to cleanly isolate the two processes, the work supplies a task-agnostic, top-down framework for diagnosing generalization dynamics that moves beyond task-specific accounts. The identification of spurious effects traceable to training choices offers a practical diagnostic contribution, and the breadth of empirical coverage across architectures strengthens the case for generality. These elements would be useful for interpretability research aimed at revealing underlying algorithms.
major comments (2)
- [Abstract and measurement sections] Abstract and measurement sections: The attribution of grokking and double descent to relative speeds rests on the claim that representational geometry, NTK spectra, and linear probing cleanly separate representation learning from readout calibration. The manuscript itself reports that non-standard recipes induce representation degradation and readout misalignment, which implies the same metrics can be sensitive to training dynamics; explicit controls or invariance tests are required to establish that the separation is not distorted by the overlapping timescales under study.
- [Grokking results] Grokking results: The statement that representation learning is gradual but not absent (contrary to lazy-to-rich) is load-bearing for the framework. Quantitative comparison of the rate of representational change versus readout change, together with ablation of the measurement tools, is needed to confirm the observed gradual shifts are not artifacts of metric contamination.
minor comments (2)
- [Methods] Clarify the precise definitions and hyper-parameters used for the representational geometry and NTK computations so that the reported signatures can be reproduced exactly.
- [Figures] Include error bars or statistical significance markers on all plots that display relative speeds or metric trajectories.
Simulated Author's Rebuttal
We thank the referee for the detailed and constructive feedback. We address the two major comments below and agree that additional controls and quantitative analyses will strengthen the manuscript.
read point-by-point responses
-
Referee: [Abstract and measurement sections] Abstract and measurement sections: The attribution of grokking and double descent to relative speeds rests on the claim that representational geometry, NTK spectra, and linear probing cleanly separate representation learning from readout calibration. The manuscript itself reports that non-standard recipes induce representation degradation and readout misalignment, which implies the same metrics can be sensitive to training dynamics; explicit controls or invariance tests are required to establish that the separation is not distorted by the overlapping timescales under study.
Authors: We agree that the potential sensitivity of the metrics warrants explicit validation. The paper already leverages this sensitivity as a diagnostic for spurious grokking and double descent under non-standard recipes. In the revision we will add invariance tests that apply the same metrics under controlled variations of training hyperparameters (while keeping the core recipe standard) to confirm that the relative-speed interpretation and separation remain stable when timescales overlap. revision: yes
-
Referee: [Grokking results] Grokking results: The statement that representation learning is gradual but not absent (contrary to lazy-to-rich) is load-bearing for the framework. Quantitative comparison of the rate of representational change versus readout change, together with ablation of the measurement tools, is needed to confirm the observed gradual shifts are not artifacts of metric contamination.
Authors: The current results already demonstrate continued representational change via multiple independent metrics (geometry, NTK, probing) that persist after readout saturation. To directly address possible contamination, the revised manuscript will include (i) explicit quantitative rate comparisons (e.g., epoch-wise slopes or integrated change) between representation and readout quantities and (ii) ablation experiments that systematically disable or replace individual measurement tools to verify that the gradual-representation conclusion is robust. revision: yes
Circularity Check
No significant circularity: empirical application of existing metrics
full rationale
The paper presents an empirical analysis that applies established tools (representational geometry, NTK spectra, linear probing) to training trajectories across tasks. No derivation chain reduces reported signatures to quantities defined by parameters fitted within the paper, nor does any central claim rest on a self-citation whose content is itself unverified or defined by the target result. The framework is self-contained against external benchmarks and falsifiable via the same measurement families on held-out runs; the reader's assessment of score 2.0 is consistent with minor non-load-bearing citations at most.
Assumptions & free parameters
assumptions (1)
- domain assumption Representation learning in the encoder and readout calibration in the final classifier can be meaningfully isolated and measured independently throughout training using representational geometry, neural tangent kernels, and linear probing.
Cite this review
Pith. "Pith review of Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent." pith.science (2026). https://pith.science/paper/IXQU5PMA
@misc{pith2026260527078,
author = {Pith},
title = {Pith review of: Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent},
year = {2026},
howpublished = {\url{https://pith.science/paper/IXQU5PMA}},
note = {Machine review of arXiv:2605.27078}
}
read the original abstract
Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena complicate this picture: in grokking, train loss falls rapidly while test performance improves abruptly only after a long delay; in epoch-wise double descent, train loss decreases monotonically while test loss or error rises and falls. Existing accounts are often task-specific, and a task-agnostic analysis framework for diagnosing and explaining these phenomena across realistic tasks and architectures is missing. We address this challenge by analyzing two competing processes that underlie learning dynamics: representation learning in the encoder and readout calibration in the final classifier. Using tools from representational geometry, neural tangent kernels, and linear probing, we show that both processes are active throughout training, with the fluctuations of their relative speed giving rise to seemingly anomalous generalization dynamics. Applying the representation-readout decomposition to grokking across a wide range of tasks and architectures, we find that the readout is train-biased before grokking onset, and representation learning is gradual but not absent, contrary to the lazy-to-rich account. The framework further provides diagnostic signatures distinguishing spurious from genuine generalization: in a previously reported MNIST grokking example and an epoch-wise double descent example, apparent delayed or non-monotone generalization is shown to arise from representation degradation and readout misalignment induced by non-standard training recipes. Together, these results establish the representation-readout decomposition as a top-down framework for understanding learning dynamics and revealing underlying algorithms for interpretability research.
Figures
Figures from the paper (32 more)
Forward citations
Cited by 2 Pith papers
-
Structure-Specific Representational Priors Causally Control the Grokking Delay
The grokking delay is causally the time to form the right feature-level representational structure, not a fixed optimization constant or a pure weight-norm effect.
-
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Muon-trained transformers grok modular addition and then lose generalization because the hidden representation and the output readout drift apart; freezing the readout and embeddings after grokking removes the collapse.
Reference graph
Works this paper leans on
-
[1]
Data augmentation consists of random crops (padding 4) and random horizontal flips; no augmentation is applied at test time
Training runs for 4,000 epochs. Data augmentation consists of random crops (padding 4) and random horizontal flips; no augmentation is applied at test time. Checkpoints are saved at 200 log-uniformly spaced epochs. Table 5: Hyperparameters for the double descent task. ResNet18 Optimizer Adam (β1, β2) (0.9,0.999) Learning rate 1×10 −4 (fixed) Weight decay ...
2023
-
[2]
Clauw et al
proposed weight sparsity as a progress measure, observing that the sparsification of weights contributing to the network output co-occurs with the onset of grokking. Clauw et al. (Clauw et al.,
-
[3]
Our framework differs from these approaches primarily in interpretability
applied O-Information—an entropy-based measure of higher-order statistical dependencies— to model activations, identifying training phases that correspond to the qualitative behavior of the train and test loss curves during grokking. Our framework differs from these approaches primarily in interpretability. While the measures described above are effective...
2023
-
[4]
and GLUE theory (see Appendix C.5), offering a robust lens that captures aspects of the transition to generalization that weight-based analyses may overlook. C.5 Previous applications of GLUE framework The Geometry Linked to Untangling Efficiency (GLUE) theory (Chou et al., 2025a), which serves as an extension of earlier work that developed Manifold Capac...
2018
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.