REVIEW 2 cited by
Progress Measures for Grokking on Real-world Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Grokking, a phenomenon where machine learning models generalize long after overfitting, has been primarily observed and studied in algorithmic tasks. This paper explores grokking in real-world datasets using deep neural networks for classification under the cross-entropy loss. We challenge the prevalent hypothesis that the $L_2$ norm of weights is the primary cause of grokking by demonstrating that grokking can occur outside the expected range of weight norms. To better understand grokking, we introduce three new progress measures: activation sparsity, absolute weight entropy, and approximate local circuit complexity. These measures are conceptually related to generalization and demonstrate a stronger correlation with grokking in real-world datasets compared to weight norms. Our findings suggest that while weight norms might usually correlate with grokking and our progress measures, they are not causative, and our proposed measures provide a better understanding of the dynamics of grokking.
Forward citations
Cited by 2 Pith papers
-
Grokking Beyond the Euclidean Norm of Model Parameters
Grokking is induced by any small nonzero regularizer whose favored solutions generalize, with a delay that scales like one over the learning rate times the regularization strength.
-
Tracing the Path to Grokking: Embeddings, Dropout, and Network Activation
The paper reports that dropout-based variance, embedding distribution shape, and neuron sparsity all shift around the moment a modular arithmetic network groks, and proposes these as forecasting signals.
Discussion (0). Continue with ORCID to comment.