Pith. sign in
Pith Number

pith:LTY3KIMI

pith:2026:LTY3KIMI745HVBQSO5UHKM37DC
not attested not anchored not stored refs pending

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

Atsuki Yamaguchi, Colin Raffel, Edward Emanuel Beeching, Elie Bakouch, Guilherme Penedo, Hynek Kydl\'i\v{c}ek, Joel Niklaus, Leandro Von Werra, Lewis Tunstall, Michal \v{S}tef\'anik, Thibaud Frere, Thomas Wolf

Rephrasing web text into structured formats like tables, FAQs, and math problems yields higher-quality synthetic pretraining data than raw web sources or prior synthetic techniques.

arxiv:2604.13977 v2 · 2026-04-15 · cs.CL · cs.AI · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{LTY3KIMI745HVBQSO5UHKM37DC}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

structured output formats, such as tables, math problems, FAQs, and tutorials, consistently outperform both curated web baselines and prior synthetic methods. Notably, increasing the size of the generator model beyond 1B parameters provides no additional benefit. By applying our findings, we develop FinePhrase, a 486-billion-token open dataset of rephrased web text that outperforms all existing synthetic data baselines while reducing generation costs by up to 30 times.

C2weakest assumption

That improvements measured in controlled experiments with smaller models and the chosen evaluation metrics will generalize to large-scale pretraining of frontier models and that the source data selection effects are not confounded by other training variables.

C3one line summary

Rephrasing web text into structured formats such as tables, math problems, FAQs, and tutorials produces higher-quality synthetic pretraining data than curated web baselines or prior synthetic methods, as demonstrated by trillion-token experiments and the resulting FinePhrase dataset that reduces gen

Receipt and verification
First computed 2026-07-31T01:33:34.789940Z
Builder pith-number-builder-2026-05-17-v1
Signature unsigned_v0
Schema pith-number/v1.0

Canonical hash

5cf1b52188ff3a7a8612776875337f18a4e2d936f5f9db01c9183f3bda6082f2

Aliases

arxiv: 2604.13977 · arxiv_version: 2604.13977v2 · doi: 10.48550/arxiv.2604.13977 · pith_short_12: LTY3KIMI745H · pith_short_16: LTY3KIMI745HVBQS · pith_short_8: LTY3KIMI
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/LTY3KIMI745HVBQSO5UHKM37DC \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 5cf1b52188ff3a7a8612776875337f18a4e2d936f5f9db01c9183f3bda6082f2
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "7c72be67364ce6a4314d5db0d044b0392f4ee56d9653b8b76b326287f8ab615f",
    "cross_cats_sorted": [
      "cs.AI",
      "cs.LG"
    ],
    "license": "http://creativecommons.org/licenses/by/4.0/",
    "primary_cat": "cs.CL",
    "submitted_at": "2026-04-15T15:24:59Z",
    "title_canon_sha256": "e0ef6b2c5f370a70f5c90870ff151fe92aab8010ef24ee73d4ec5b622abc94b2"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2604.13977",
    "kind": "arxiv",
    "version": 2
  }
}