pith:LTY3KIMI
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
Rephrasing web text into structured formats like tables, FAQs, and math problems yields higher-quality synthetic pretraining data than raw web sources or prior synthetic techniques.
arxiv:2604.13977 v2 · 2026-04-15 · cs.CL · cs.AI · cs.LG
Add to your LaTeX paper
\usepackage{pith}
\pithnumber{LTY3KIMI745HVBQSO5UHKM37DC}
Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge
Record completeness
Claims
structured output formats, such as tables, math problems, FAQs, and tutorials, consistently outperform both curated web baselines and prior synthetic methods. Notably, increasing the size of the generator model beyond 1B parameters provides no additional benefit. By applying our findings, we develop FinePhrase, a 486-billion-token open dataset of rephrased web text that outperforms all existing synthetic data baselines while reducing generation costs by up to 30 times.
That improvements measured in controlled experiments with smaller models and the chosen evaluation metrics will generalize to large-scale pretraining of frontier models and that the source data selection effects are not confounded by other training variables.
Rephrasing web text into structured formats such as tables, math problems, FAQs, and tutorials produces higher-quality synthetic pretraining data than curated web baselines or prior synthetic methods, as demonstrated by trillion-token experiments and the resulting FinePhrase dataset that reduces gen
Receipt and verification
| First computed | 2026-07-31T01:33:34.789940Z |
|---|---|
| Builder | pith-number-builder-2026-05-17-v1 |
| Signature | unsigned_v0 |
| Schema | pith-number/v1.0 |
Canonical hash
5cf1b52188ff3a7a8612776875337f18a4e2d936f5f9db01c9183f3bda6082f2
Aliases
· · · · ·Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/LTY3KIMI745HVBQSO5UHKM37DC \
| jq -c '.canonical_record' \
| python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 5cf1b52188ff3a7a8612776875337f18a4e2d936f5f9db01c9183f3bda6082f2
Canonical record JSON
{
"metadata": {
"abstract_canon_sha256": "7c72be67364ce6a4314d5db0d044b0392f4ee56d9653b8b76b326287f8ab615f",
"cross_cats_sorted": [
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"primary_cat": "cs.CL",
"submitted_at": "2026-04-15T15:24:59Z",
"title_canon_sha256": "e0ef6b2c5f370a70f5c90870ff151fe92aab8010ef24ee73d4ec5b622abc94b2"
},
"schema_version": "1.0",
"source": {
"id": "2604.13977",
"kind": "arxiv",
"version": 2
}
}