Pith. sign in

REVIEW 2 cited by

Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17247 v2 pith:LX6JF7ZZ submitted 2023-10-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords grokkingnetworksneuralsettingscomplexityempiricalmodelphenomenon
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In some settings neural networks exhibit a phenomenon known as \textit{grokking}, where they achieve perfect or near-perfect accuracy on the validation set long after the same performance has been achieved on the training set. In this paper, we discover that grokking is not limited to neural networks but occurs in other settings such as Gaussian process (GP) classification, GP regression, linear regression and Bayesian neural networks. We also uncover a mechanism by which to induce grokking on algorithmic datasets via the addition of dimensions containing spurious information. The presence of the phenomenon in non-neural architectures shows that grokking is not restricted to settings considered in current theoretical and empirical studies. Instead, grokking may be possible in any model where solution search is guided by complexity and error.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model

    cs.LG 2025-04 conditional novelty 6.0 of 10

    GrokTransfer transfers an embedding learned by a small 'weaker' model to a larger model, eliminating the grokking delay so the target model generalizes almost immediately.

  2. Not All Explanations for Deep Learning Phenomena Are Equally Valuable

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A position paper arguing that narrow, puzzle-solving explanations of deep learning edge case phenomena are low-value, and that these phenomena should instead be used to stress-test broad explanatory theories.

Pith tools