Pith. sign in

REVIEW 1 cited by

Grokking as Compression: A Nonlinear Complexity Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05918 v1 pith:3V7SYW4G submitted 2023-10-09 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords complexitycompressionlinearnetworkgeneralizationgrokkingneuralnumber
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We attribute grokking, the phenomenon where generalization is much delayed after memorization, to compression. To do so, we define linear mapping number (LMN) to measure network complexity, which is a generalized version of linear region number for ReLU networks. LMN can nicely characterize neural network compression before generalization. Although the $L_2$ norm has been a popular choice for characterizing model complexity, we argue in favor of LMN for a number of reasons: (1) LMN can be naturally interpreted as information/computation, while $L_2$ cannot. (2) In the compression phase, LMN has linear relations with test losses, while $L_2$ is correlated with test losses in a complicated nonlinear way. (3) LMN also reveals an intriguing phenomenon of the XOR network switching between two generalization solutions, while $L_2$ does not. Besides explaining grokking, we argue that LMN is a promising candidate as the neural network version of the Kolmogorov complexity since it explicitly considers local or conditioned linear computations aligned with the nature of modern artificial neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Do Neural Networks Learn World Models?

    cs.LG 2025-02 conditional novelty 7.0 of 10

    With Boolean variables, a low-degree bias, and a task distribution weighted toward simple functions of the latents, multi-task training provably recovers the latent world model up to permutations and negations.

Pith tools