Pith. sign in

REVIEW 3 cited by

Progress Measures for Grokking on Real-world Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12755 v2 pith:VVNFYRGL submitted 2024-05-21 cs.LG cs.AI

Progress Measures for Grokking on Real-world Tasks

classification cs.LG cs.AI
keywords grokkingmeasuresweightnormsprogressreal-worldbetterdatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Grokking, a phenomenon where machine learning models generalize long after overfitting, has been primarily observed and studied in algorithmic tasks. This paper explores grokking in real-world datasets using deep neural networks for classification under the cross-entropy loss. We challenge the prevalent hypothesis that the $L_2$ norm of weights is the primary cause of grokking by demonstrating that grokking can occur outside the expected range of weight norms. To better understand grokking, we introduce three new progress measures: activation sparsity, absolute weight entropy, and approximate local circuit complexity. These measures are conceptually related to generalization and demonstrate a stronger correlation with grokking in real-world datasets compared to weight norms. Our findings suggest that while weight norms might usually correlate with grokking and our progress measures, they are not causative, and our proposed measures provide a better understanding of the dynamics of grokking.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

    cs.LG 2026-05 unverdicted novelty 7.0

    Slingshot loss spikes result from floating-point precision limits that round correct-class gradients to zero, triggering Numerical Feature Inflation and breaking gradient zero-sum constraints.

  2. Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

    cs.LG 2026-05 unverdicted novelty 7.0

    Slingshot loss spikes arise from floating-point precision limits that round correct-class gradients to zero, breaking zero-sum constraints and driving exponential parameter growth through numerical feature inflation.

  3. Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

    cs.LG 2026-05 unverdicted novelty 7.0

    Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.