REVIEW 1 cited by
Controlling Grokking with Nonlinearity and Data Symmetry
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper demonstrates that grokking behavior in modular arithmetic with a modulus P in a neural network can be controlled by modifying the profile of the activation function as well as the depth and width of the model. Plotting the even PCA projections of the weights of the last NN layer against their odd projections further yields patterns which become significantly more uniform when the nonlinearity is increased by incrementing the number of layers. These patterns can be employed to factor P when P is nonprime. Finally, a metric for the generalization ability of the network is inferred from the entropy of the layer weights while the degree of nonlinearity is related to correlations between the local entropy of the weights of the neurons in the final layer.
Forward citations
Cited by 1 Pith paper
-
Tracing the Path to Grokking: Embeddings, Dropout, and Network Activation
The paper reports that dropout-based variance, embedding distribution shape, and neuron sparsity all shift around the moment a modular arithmetic network groks, and proposes these as forecasting signals.
Discussion (0). Continue with ORCID to comment.