FSD, a permutation-tested metric of Fourier circuit synchronization, precedes grokking by a mean of 1722 steps across nine modular addition setups and causally controls grokking timing when weight decay is varied at the FSD ceiling.
arXiv preprint arXiv:2405.20233 , year=
7 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 7representative citing papers
Normal alignment is the rank-one Jacobian structure that lets classifiers minimize loss and maximize local robustness in sparse regimes; the paper proves its optimality and uses it to create GrokAlign and RFAMs.
ILDR detects the geometric reorganization preceding grokking by measuring when inter-class centroid separation exceeds intra-class scatter by 2.5 times its baseline in penultimate-layer representations.
EGD equalizes gradient speeds across singular directions, eliminating or shortening grokking plateaus on modular addition and sparse parity problems.
A norm penalty constraining activations to a sqrt(d)-radius hypersphere accelerates grokking by up to 6x on modular arithmetic via radial suppression in activation dynamics.
Random Matrix Theory detects overfitting via growing Correlation Traps in weight spectra during the anti-grokking phase of neural network training.
Grokking emerges near the model size where memorization timescale T_mem(P) intersects generalization timescale T_gen(P) on modular arithmetic.
citing papers explorer
-
Circuit Synchronization Precedes Generalization: A Causal Precursor to Grokking
FSD, a permutation-tested metric of Fourier circuit synchronization, precedes grokking by a mean of 1722 steps across nine modular addition setups and causally controls grokking timing when weight decay is varied at the FSD ceiling.
-
The Geometric Structure of Models Learning Sparse Data
Normal alignment is the rank-one Jacobian structure that lets classifiers minimize loss and maximize local robustness in sparse regimes; the paper proves its optimality and uses it to create GrokAlign and RFAMs.
-
ILDR: Geometric Early Detection of Grokking
ILDR detects the geometric reorganization preceding grokking by measuring when inter-class centroid separation exceeds intra-class scatter by 2.5 times its baseline in penultimate-layer representations.
-
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
EGD equalizes gradient speeds across singular directions, eliminating or shortening grokking plateaus on modular addition and sparse parity problems.
-
Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization
A norm penalty constraining activations to a sqrt(d)-radius hypersphere accelerates grokking by up to 6x on modular arithmetic via radial suppression in activation dynamics.
-
Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory
Random Matrix Theory detects overfitting via growing Correlation Traps in weight spectra during the anti-grokking phase of neural network training.
-
Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds
Grokking emerges near the model size where memorization timescale T_mem(P) intersects generalization timescale T_gen(P) on modular arithmetic.