Pith. sign in

REVIEW 1 cited by

An Information Theory of Compute-Optimal Size Scaling, Emergence, and Plateaus in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01243 v2 pith:CRA5KNN6 submitted 2024-10-02 cs.IT math.IT

classification cs.ITmath.IT
keywords languagescalingsizemodelscompute-optimaltheorydecodingemergence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent empirical studies show three phenomena with increasing size of language models: compute-optimal size scaling, emergent capabilities, and performance plateauing. We present a simple unified mathematical framework to explain all of these language model scaling phenomena, building on recent skill-text bipartite graph frameworks for semantic learning. Modeling the learning of concepts from texts as an iterative process yields an analogy to iterative decoding of low-density parity check (LDPC) codes in information theory. Thence, drawing on finite-size scaling characterizations of LDPC decoding, we derive the compute-optimal size scaling (Chinchilla rule) for language models. Further, using tools from random network theory, we provide a simple explanation for both emergence of complex skills and plateauing of performance as the size of language models scale. We see multiple plateaus.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLOUD: A Scalable and Physics-Informed Foundation Model for Crystal Representation Learning

    cond-mat.mtrl-sci 2025-06 conditional novelty 6.0 of 10

    CLOUD, a BERT-style model pretrained on 6.3 million crystal structures with a new symmetry-aware string encoding (SCOPE), gives competitive property predictions and, when combined with the Debye model, extrapolates he...

Pith tools