pith:6FLJ2FRK
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Larger language models up to 280 billion parameters reach state-of-the-art results on most of 152 tasks, with scale helping reading and fact-checking most.
arxiv:2112.11446 v2 · 2021-12-08 · cs.CL · cs.AI
Add to your LaTeX paper
\usepackage{pith}
\pithnumber{6FLJ2FRK3565SJWTG3QRIA4TF7}
Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge
Record completeness
Claims
These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, and the identification of toxic language, but logical and mathematical reasoning see less benefit.
That performance differences across scales are primarily driven by model size rather than confounding factors such as dataset composition, training details, or evaluation choices, and that the 152 tasks sufficiently represent broader capabilities.
Gopher, a 280 billion parameter language model, achieves state-of-the-art performance on the majority of 152 tasks with largest gains in reading comprehension, fact-checking, and toxic language detection.
References
Cited by
Receipt and verification
| First computed | 2026-07-05T03:50:27.025729Z |
|---|---|
| Builder | pith-number-builder-2026-05-17-v1 |
| Signature | Pith Ed25519
(pith-v1-2026-05) · public key |
| Schema | pith-number/v1.0 |
Canonical hash
f1569d162adf7dd926d336e11403932fc838cd859c5f0bd7573161321d5279e0
Aliases
· · · · ·Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7 \
| jq -c '.canonical_record' \
| python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: f1569d162adf7dd926d336e11403932fc838cd859c5f0bd7573161321d5279e0
Canonical record JSON
{
"metadata": {
"abstract_canon_sha256": "acb60e61a0467c59a348897bf95d6b42216d5834406a43f0c68ae829e874e70c",
"cross_cats_sorted": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"primary_cat": "cs.CL",
"submitted_at": "2021-12-08T19:41:47Z",
"title_canon_sha256": "3398a07379e6d2893eae299e46e16b8b3997b3d623cc8b591e7e13c1ce02542d"
},
"schema_version": "1.0",
"source": {
"id": "2112.11446",
"kind": "arxiv",
"version": 2
}
}