REVIEW 5 cited by
Grokking as a First Order Phase Transition in Two Layer Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A key property of deep neural networks (DNNs) is their ability to learn new features during training. This intriguing aspect of deep learning stands out most clearly in recently reported Grokking phenomena. While mainly reflected as a sudden increase in test accuracy, Grokking is also believed to be a beyond lazy-learning/Gaussian Process (GP) phenomenon involving feature learning. Here we apply a recent development in the theory of feature learning, the adaptive kernel approach, to two teacher-student models with cubic-polynomial and modular addition teachers. We provide analytical predictions on feature learning and Grokking properties of these models and demonstrate a mapping between Grokking and the theory of phase transitions. We show that after Grokking, the state of the DNN is analogous to the mixed phase following a first-order phase transition. In this mixed phase, the DNN generates useful internal representations of the teacher that are sharply distinct from those before the transition.
Forward citations
Cited by 5 Pith papers
-
Learning words in groups: fusion algebras, tensor ranks and grokking
Group word operations can be learned by small two-layer networks because the associated word tensor has low rank, decomposable through the fusion algebra of the group's self-conjugate representations.
-
Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations
A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.
-
Grokking vs. Learning: Same Features, Different Encodings
Grokked and steadily trained models learn the same features, but steady training can produce much more compressible models in a parameter regime that grokking does not reach.
-
Consciousness as a Jamming Phase
Large language models are claimed to become conscious when their word embeddings jam into a critically correlated state, but no evidence is provided.
-
Rethinking Over-Smoothing in Graph Neural Networks: A Perspective from Anderson Localization
A single-author preprint re-frames GNN over-smoothing as Anderson localization, defining a participation-degree metric and proposing degree-dependent edge reweighting as mitigation, without proof or experiments.
Discussion (0). Continue with ORCID to comment.