REVIEW 3 cited by
Predicting Grokking Long Before it Happens: A look into the loss landscape of models which grok
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper focuses on predicting the occurrence of grokking in neural networks, a phenomenon in which perfect generalization emerges long after signs of overfitting or memorization are observed. It has been reported that grokking can only be observed with certain hyper-parameters. This makes it critical to identify the parameters that lead to grokking. However, since grokking occurs after a large number of epochs, searching for the hyper-parameters that lead to it is time-consuming. In this paper, we propose a low-cost method to predict grokking without training for a large number of epochs. In essence, by studying the learning curve of the first few epochs, we show that one can predict whether grokking will occur later on. Specifically, if certain oscillations occur in the early epochs, one can expect grokking to occur if the model is trained for a much longer period of time. We propose using the spectral signature of a learning curve derived by applying the Fourier transform to quantify the amplitude of low-frequency components to detect the presence of such oscillations. We also present additional experiments aimed at explaining the cause of these oscillations and characterizing the loss landscape.
Forward citations
Cited by 3 Pith papers
-
To Grok Grokking: Provable Grokking in Ridge Regression
Training loss in over-parameterized ridge regression drops fast along data directions, while test loss waits for slow weight-decay shrinkage of the perpendicular directions, yielding provable grokking with delay ~1/λ.
-
Grokking Beyond the Euclidean Norm of Model Parameters
Grokking is induced by any small nonzero regularizer whose favored solutions generalize, with a delay that scales like one over the learning rate times the regularization strength.
-
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
CRAFT combines ConceptNet or LLM-derived affordance priors with CLIP visual similarity in an iterative energy-based re-ranking loop, achieving the best non-oracle accuracy on a functional affordance grounding benchmark.
Discussion (0). Continue with ORCID to comment.