Pith. sign in

REVIEW 2 cited by

Pretraining on the Test Set Is All You Need

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08632 v1 pith:YOM5ZNLW submitted 2023-09-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords benchmarksdataevaluationmixturemodelsnovelphi-ctnltextbf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inspired by recent work demonstrating the promise of smaller Transformer-based language models pretrained on carefully curated data, we supercharge such approaches by investing heavily in curating a novel, high quality, non-synthetic data mixture based solely on evaluation benchmarks. Using our novel dataset mixture consisting of less than 100 thousand tokens, we pretrain a 1 million parameter transformer-based LLM \textbf{phi-CTNL} (pronounced ``fictional") that achieves perfect results across diverse academic benchmarks, strictly outperforming all known foundation models. \textbf{phi-CTNL} also beats power-law scaling and exhibits a never-before-seen grokking-like ability to accurately predict downstream evaluation benchmarks' canaries.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Models Improve When Pretraining Data Matches Target Tasks

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.

  2. Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A debate-based evaluation protocol on 50 MMLU-Pro questions: fine-tuning on the test set boosts standard accuracy from 50% to 82% but not debate win rates.

Pith tools