Pith. sign in

REVIEW 5 cited by

Deep Networks Always Grok and Here is Why

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15555 v2 pith:JKFCN56O submitted 2024-02-23 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords trainingdelayedgeneralizationgrokkingmappingregionscomplexitydeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Grokking, or delayed generalization, is a phenomenon where generalization in a deep neural network (DNN) occurs long after achieving near zero training error. Previous studies have reported the occurrence of grokking in specific controlled settings, such as DNNs initialized with large-norm parameters or transformers trained on algorithmic datasets. We demonstrate that grokking is actually much more widespread and materializes in a wide range of practical settings, such as training of a convolutional neural network (CNN) on CIFAR10 or a Resnet on Imagenette. We introduce the new concept of delayed robustness, whereby a DNN groks adversarial examples and becomes robust, long after interpolation and/or generalization. We develop an analytical explanation for the emergence of both delayed generalization and delayed robustness based on the local complexity of a DNN's input-output mapping. Our local complexity measures the density of so-called linear regions (aka, spline partition regions) that tile the DNN input space and serves as a utile progress measure for training. We provide the first evidence that, for classification problems, the linear regions undergo a phase transition during training whereafter they migrate away from the training samples (making the DNN mapping smoother there) and towards the decision boundary (making the DNN mapping less smooth there). Grokking occurs post phase transition as a robust partition of the input space thanks to the linearization of the DNN mapping around the training points. Website: https://bit.ly/grok-adversarial

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning words in groups: fusion algebras, tensor ranks and grokking

    cs.LG 2025-09 conditional novelty 8.0 of 10

    Group word operations can be learned by small two-layer networks because the associated word tensor has low rank, decomposable through the fusion algebra of the group's self-conjugate representations.

  2. Emergent Generalization by Representation Learning in Artificial Neural Networks

    q-bio.NC 2026-07 conditional novelty 6.0 of 10

    An explicit low-dimensional bottleneck is necessary for OOD generalisation in reservoir networks, and the non-monotonic rise of causal emergence in the latent code predicts generalisation both in silico and in mouse CA1.

  3. The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Post-memorization learning in grokking is equivalent to minimizing the weight norm on the zero-loss manifold, with a closed-form approximation for two-layer networks.

  4. Grokking vs. Learning: Same Features, Different Encodings

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Grokked and steadily trained models learn the same features, but steady training can produce much more compressible models in a parameter regime that grokking does not reach.

  5. Mechanistic Insights into Grokking from the Embedding Layer

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Trainable embeddings in a simple MLP cause delayed generalization (grokking) on modular arithmetic, and a higher embedding learning rate plus balanced sampling accelerates it.

Pith tools