Pith. sign in

REVIEW 2 cited by

Progress Measures for Grokking on Real-world Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12755 v2 pith:VVNFYRGL submitted 2024-05-21 cs.LG cs.AI

classification cs.LGcs.AI
keywords grokkingmeasuresweightnormsprogressreal-worldbetterdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Grokking, a phenomenon where machine learning models generalize long after overfitting, has been primarily observed and studied in algorithmic tasks. This paper explores grokking in real-world datasets using deep neural networks for classification under the cross-entropy loss. We challenge the prevalent hypothesis that the $L_2$ norm of weights is the primary cause of grokking by demonstrating that grokking can occur outside the expected range of weight norms. To better understand grokking, we introduce three new progress measures: activation sparsity, absolute weight entropy, and approximate local circuit complexity. These measures are conceptually related to generalization and demonstrate a stronger correlation with grokking in real-world datasets compared to weight norms. Our findings suggest that while weight norms might usually correlate with grokking and our progress measures, they are not causative, and our proposed measures provide a better understanding of the dynamics of grokking.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grokking Beyond the Euclidean Norm of Model Parameters

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Grokking is induced by any small nonzero regularizer whose favored solutions generalize, with a delay that scales like one over the learning rate times the regularization strength.

  2. Tracing the Path to Grokking: Embeddings, Dropout, and Network Activation

    cs.LG 2025-07 reject novelty 4.0 of 10

    The paper reports that dropout-based variance, embedding distribution shape, and neuron sparsity all shift around the moment a modular arithmetic network groks, and proposes these as forecasting signals.

Pith tools