Towards Understanding Grokking: An Effective Theory of Representation Learning

Eric J. Michaud; Max Tegmark; Mike Williams; Niklas Nolte; Ouail Kitouni; Ziming Liu

arxiv: 2205.10343 · v2 · pith:BKSQ5JNDnew · submitted 2022-05-20 · 💻 cs.LG · cond-mat.dis-nn· cond-mat.stat-mech· cs.AI· physics.class-ph

Towards Understanding Grokking: An Effective Theory of Representation Learning

Ziming Liu , Ouail Kitouni , Niklas Nolte , Eric J. Michaud , Max Tegmark , Mike Williams This is my paper

classification 💻 cs.LG cond-mat.dis-nncond-mat.stat-mechcs.AIphysics.class-ph

keywords grokkingphaselearningeffectivecomprehensionfindmemorizationtheory

0 comments

read the original abstract

We aim to understand grokking, a phenomenon where models generalize long after overfitting their training set. We present both a microscopic analysis anchored by an effective theory and a macroscopic analysis of phase diagrams describing learning performance across hyperparameters. We find that generalization originates from structured representations whose training dynamics and dependence on training set size can be predicted by our effective theory in a toy setting. We observe empirically the presence of four learning phases: comprehension, grokking, memorization, and confusion. We find representation learning to occur only in a "Goldilocks zone" (including comprehension and grokking) between memorization and confusion. We find on transformers the grokking phase stays closer to the memorization phase (compared to the comprehension phase), leading to delayed generalization. The Goldilocks phase is reminiscent of "intelligence from starvation" in Darwinian evolution, where resource limitations drive discovery of more efficient solutions. This study not only provides intuitive explanations of the origin of grokking, but also highlights the usefulness of physics-inspired tools, e.g., effective theories and phase diagrams, for understanding deep learning.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Progress measures for grokking via mechanistic interpretability
cs.LG 2023-01 accept novelty 8.0

Grokking arises from gradual amplification of a Fourier-based circuit in the weights followed by removal of memorizing components.
The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior
cs.LG 2026-03 unverdicted novelty 7.0

The grokking delay in encoder-decoder models on one-step Collatz prediction stems from decoder inability to use early-learned encoder representations of parity and residue structure, with numeral base acting as a stro...
Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory
cs.LG 2026-05 unverdicted novelty 6.0

A Random Matrix Theory method identifies growing Correlation Traps in neural network weight spectra during an 'anti-grokking' overfitting phase, and applies the same diagnostic to some foundation LLMs.
Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory
cs.LG 2026-05 unverdicted novelty 6.0

Random Matrix Theory detects overfitting via growing Correlation Traps in weight spectra during the anti-grokking phase of neural network training.
Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry
cs.LG 2026-05 unverdicted novelty 5.0

Proposes a two-gradient-field model with candidate order parameters alpha_dagger and kappa_c to unify phase transitions across learning theory and non-equilibrium chemistry.