Pith. sign in

REVIEW 4 cited by

Algorithmic progress in language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05812 v1 pith:XYD6E64R submitted 2024-03-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords languageprogressalgorithmicalgorithmscomputemodelsanalysiscontributions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 95% confidence interval of around 5 to 14 months, substantially faster than hardware gains per Moore's Law. We estimate augmented scaling laws, which enable us to quantify algorithmic progress and determine the relative contributions of scaling models versus innovations in training algorithms. Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period. Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. DICE: Data Influence Cascade in Decentralized Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    DICE defines and approximates multi-hop data influence in decentralized learning, showing that influence is shaped by data, topology, and loss curvature.

  2. A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.

  3. Technical Requirements for Halting Dangerous AI Activities

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A taxonomy of compute-centric technical interventions, graded by readiness and mapped to five AI governance plans, argues that halting dangerous AI requires substantial control over AI compute.

  4. Meek Models Shall Inherit the Earth

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Under fixed-distribution neural scaling laws, the capability gap between state-of-the-art and low-compute AI models shrinks over time toward zero.

Pith tools