Pith. sign in

REVIEW 1 cited by

Critical Data Size of Language Models from a Grokking Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.10463 v3 pith:XMHGTEVU submitted 2024-01-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagedatamodelscriticalgrokkingsizeconfigurationefficiency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore the critical data size in language models, a threshold that marks a fundamental shift from quick memorization to slow generalization. We formalize the phase transition under the grokking configuration into the Data Efficiency Hypothesis and identify data insufficiency, sufficiency, and surplus regimes in language models training dynamics. We develop a grokking configuration to reproduce grokking on simplistic language models stably by rescaling initialization and weight decay. We show that generalization occurs only when language models reach a critical size. We analyze grokking across sample-wise and model-wise, verifying the proposed data efficiency hypothesis. Our experiments reveal smoother phase transitions occurring at the critical dataset size for language datasets. As the model size increases, this critical point also becomes larger, indicating that larger models require more data. Our results deepen the understanding of language model training, offering a novel perspective on the role of data in the learning mechanism of language models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hybrid Reasoning for Perception, Explanation, and Autonomous Action in Manufacturing

    cs.AI 2025-06 conditional novelty 6.0 of 10

    CIPHER embeds a convolutional regression expert into a vision-language-action model, enabling a commercial 3D printer to perceive extrusion state, reason about faults, and generate corrective G-code from images or text.

Pith tools