{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:6FLJ2FRK3565SJWTG3QRIA4TF7","short_pith_number":"pith:6FLJ2FRK","schema_version":"1.0","canonical_sha256":"f1569d162adf7dd926d336e11403932fc838cd859c5f0bd7573161321d5279e0","source":{"kind":"arxiv","id":"2112.11446","version":2},"attestation_state":"computed","paper":{"title":"Scaling Language Models: Methods, Analysis & Insights from Training Gopher","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Larger language models up to 280 billion parameters reach state-of-the-art results on most of 152 tasks, with scale helping reading and fact-checking most.","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Adhiguna Kuncoro, Aidan Clark, Aida Nematzadeh, Albin Cassirer, Amelia Glaese, Amy Wu, Angeliki Lazaridou, Antonia Creswell, Arthur Mensch, Aurelia Guy, Blake Hechtman, Chris Dyer, Chris Jones, Cyprien de Masson d'Autume, Daniel Toyama, David Budden, Demis Hassabis, Diego de las Casas, Domenic Donato, Doug Fritz, Ed Lockhart, Elena Buchatskaya, Elena Gribovskaya, Eliza Rutherford, Erich Elsen, Esme Sutherland, Francis Song, Geoffrey Irving, George van den Driessche, Iason Gabriel, Igor Babuschkin, Irina Higgins, Jack W. Rae, Jacob Menick, James Bradbury, Jean-Baptiste Lespiau, Jeff Stanway, Johannes Welbl, John Aslanides, John Mellor, Jonathan Uesato, Jordan Hoffmann, Kareem Ayoub, Karen Simonyan, Katie Millican, Koray Kavukcuoglu, Laura Rimell, Laura Weidinger, Laurent Sifre, Lena Martens, Lisa Anne Hendricks, Lorrayne Bennett, Mantas Pajarskas, Maria Tsimpoukelli, Maribeth Rauh, Matthew Johnson, Michela Paganini, Nat McAleese, Nikolai Grigorev, Oriol Vinyals, Po-Sen Huang, Richard Powell, Roman Ring, Saffron Huang, Sarah Henderson, Sebastian Borgeaud, Siddhant Jayakumar, Simon Osindero, Sumanth Dathathri, Susannah Young, Tayfun Terzi, Thibault Sottiaux, Toby Pohlen, Tom Hennigan, Trevor Cai, Vladimir Mikulik, William Isaac, Xiang Lorraine Li, Yujia Li, Zhitao Gong","submitted_at":"2021-12-08T19:41:47Z","abstract_excerpt":"Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model performance across a wide range of model scales -- from models with tens of millions of parameters up to a 280 billion parameter model called Gopher. These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, an"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":false},"canonical_record":{"source":{"id":"2112.11446","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2021-12-08T19:41:47Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"3398a07379e6d2893eae299e46e16b8b3997b3d623cc8b591e7e13c1ce02542d","abstract_canon_sha256":"acb60e61a0467c59a348897bf95d6b42216d5834406a43f0c68ae829e874e70c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T03:50:27.026274Z","signature_b64":"uvcSkFuIaFwFfu9+FF8B+9kH+X7pcYmYRVzXx657B/Tj4kllDpKLQqw52aez67Sj09BGa3Sp3F4VS4DYuFmcDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"f1569d162adf7dd926d336e11403932fc838cd859c5f0bd7573161321d5279e0","last_reissued_at":"2026-07-05T03:50:27.025729Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T03:50:27.025729Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Scaling Language Models: Methods, Analysis & Insights from Training Gopher","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Larger language models up to 280 billion parameters reach state-of-the-art results on most of 152 tasks, with scale helping reading and fact-checking most.","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Adhiguna Kuncoro, Aidan Clark, Aida Nematzadeh, Albin Cassirer, Amelia Glaese, Amy Wu, Angeliki Lazaridou, Antonia Creswell, Arthur Mensch, Aurelia Guy, Blake Hechtman, Chris Dyer, Chris Jones, Cyprien de Masson d'Autume, Daniel Toyama, David Budden, Demis Hassabis, Diego de las Casas, Domenic Donato, Doug Fritz, Ed Lockhart, Elena Buchatskaya, Elena Gribovskaya, Eliza Rutherford, Erich Elsen, Esme Sutherland, Francis Song, Geoffrey Irving, George van den Driessche, Iason Gabriel, Igor Babuschkin, Irina Higgins, Jack W. Rae, Jacob Menick, James Bradbury, Jean-Baptiste Lespiau, Jeff Stanway, Johannes Welbl, John Aslanides, John Mellor, Jonathan Uesato, Jordan Hoffmann, Kareem Ayoub, Karen Simonyan, Katie Millican, Koray Kavukcuoglu, Laura Rimell, Laura Weidinger, Laurent Sifre, Lena Martens, Lisa Anne Hendricks, Lorrayne Bennett, Mantas Pajarskas, Maria Tsimpoukelli, Maribeth Rauh, Matthew Johnson, Michela Paganini, Nat McAleese, Nikolai Grigorev, Oriol Vinyals, Po-Sen Huang, Richard Powell, Roman Ring, Saffron Huang, Sarah Henderson, Sebastian Borgeaud, Siddhant Jayakumar, Simon Osindero, Sumanth Dathathri, Susannah Young, Tayfun Terzi, Thibault Sottiaux, Toby Pohlen, Tom Hennigan, Trevor Cai, Vladimir Mikulik, William Isaac, Xiang Lorraine Li, Yujia Li, Zhitao Gong","submitted_at":"2021-12-08T19:41:47Z","abstract_excerpt":"Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model performance across a wide range of model scales -- from models with tens of millions of parameters up to a 280 billion parameter model called Gopher. These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, an"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, and the identification of toxic language, but logical and mathematical reasoning see less benefit.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That performance differences across scales are primarily driven by model size rather than confounding factors such as dataset composition, training details, or evaluation choices, and that the 152 tasks sufficiently represent broader capabilities.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Gopher, a 280 billion parameter language model, achieves state-of-the-art performance on the majority of 152 tasks with largest gains in reading comprehension, fact-checking, and toxic language detection.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Larger language models up to 280 billion parameters reach state-of-the-art results on most of 152 tasks, with scale helping reading and fact-checking most.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"55bfbcfd91deda931d1e3aca0d988cf1257c64db2f7fe50a207729769354f479"},"source":{"id":"2112.11446","kind":"arxiv","version":2},"verdict":{"id":"ee88b084-9d70-45a1-a804-6c644045bea1","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T19:07:09.329416Z","strongest_claim":"These models are evaluated on 152 diverse tasks, achieving state-of-the-art performance across the majority. Gains from scale are largest in areas such as reading comprehension, fact-checking, and the identification of toxic language, but logical and mathematical reasoning see less benefit.","one_line_summary":"Gopher, a 280 billion parameter language model, achieves state-of-the-art performance on the majority of 152 tasks with largest gains in reading comprehension, fact-checking, and toxic language detection.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That performance differences across scales are primarily driven by model size rather than confounding factors such as dataset composition, training details, or evaluation choices, and that the 152 tasks sufficiently represent broader capabilities.","pith_extraction_headline":"Larger language models up to 280 billion parameters reach state-of-the-art results on most of 152 tasks, with scale helping reading and fact-checking most."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2112.11446/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":79,"sample":[{"doi":"","year":2020,"title":"URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb49674 18bfb8ac142f64a-Paper.pdf. J. Buckman. Fair ML tools require problematic ML models.https://jacobbuckman.com/2021- 02-15-fair-ml-too","work_id":"0ceeaa26-b5a4-4dee-9e71-38b724ab84af","ref_index":1,"cited_arxiv_id":"1812.01193","is_internal_anchor":true},{"doi":"10.18653/v1/2020.findings-emnlp.301","year":1902,"title":"doi: 10.18653/v1/2020.findings-emnlp.301","work_id":"b0252c97-1b16-43d3-9417-9f8a72176a9f","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.1145/3351095.3372826","year":2009,"title":", author Barocas, S","work_id":"f306132e-972e-4117-a0f7-ce2b4081c431","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.18653/v1/2020.findings-emnlp.372","year":2011,"title":"In: Cohn, T., He, Y., Liu, Y","work_id":"d178d322-4363-48e5-bede-e7e4ebf4a434","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.1145/3360307","year":2016,"title":"Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David A","work_id":"139ebadb-ddc7-49a8-a723-31b8d3d2caed","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":79,"snapshot_sha256":"1b7abe520bee470b6920c62904145c36cea3dc75c7e5c78e28d1b27bba5609a5","internal_anchors":2},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2112.11446","created_at":"2026-07-05T03:50:27.025790+00:00"},{"alias_kind":"arxiv_version","alias_value":"2112.11446v2","created_at":"2026-07-05T03:50:27.025790+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2112.11446","created_at":"2026-07-05T03:50:27.025790+00:00"},{"alias_kind":"pith_short_12","alias_value":"6FLJ2FRK3565","created_at":"2026-07-05T03:50:27.025790+00:00"},{"alias_kind":"pith_short_16","alias_value":"6FLJ2FRK3565SJWT","created_at":"2026-07-05T03:50:27.025790+00:00"},{"alias_kind":"pith_short_8","alias_value":"6FLJ2FRK","created_at":"2026-07-05T03:50:27.025790+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":119,"internal_anchor_count":119,"sample":[{"citing_arxiv_id":"2606.25008","citing_title":"Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22473","citing_title":"Interleaved Speech Language Models Latently Work In Text","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19712","citing_title":"Efficient Neural Network Model Selection for Few-Class Application Datasets","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19468","citing_title":"Characterizing Narrative Content in Web-scale LLM Pretraining Data","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18650","citing_title":"BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training","ref_index":33,"is_internal_anchor":true},{"citing_arxiv_id":"2607.02447","citing_title":"Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12234","citing_title":"On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study","ref_index":34,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11499","citing_title":"Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality","ref_index":10,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10554","citing_title":"Benchmarking Knowledge Editing using Logical Rules","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10706","citing_title":"Unifying Data, Memory, and Compute Efficiency in LLM training: A Survey","ref_index":115,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07019","citing_title":"PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer","ref_index":43,"is_internal_anchor":true},{"citing_arxiv_id":"2606.04661","citing_title":"CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts","ref_index":182,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02054","citing_title":"eMoT: evolving Memory-of-Thought via Symbolic Anchoring and Memory Corrosion","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2606.30775","citing_title":"A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"2605.23660","citing_title":"Using Large Language Models in Physics Education","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2605.29548","citing_title":"Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07597","citing_title":"Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them","ref_index":69,"is_internal_anchor":true},{"citing_arxiv_id":"2605.23660","citing_title":"Using Large Language Models in Physics Education","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2204.06745","citing_title":"GPT-NeoX-20B: An Open-Source Autoregressive Language Model","ref_index":73,"is_internal_anchor":true},{"citing_arxiv_id":"2301.13688","citing_title":"The Flan Collection: Designing Data and Methods for Effective Instruction Tuning","ref_index":45,"is_internal_anchor":true},{"citing_arxiv_id":"2312.11805","citing_title":"Gemini: A Family of Highly Capable Multimodal Models","ref_index":79,"is_internal_anchor":true},{"citing_arxiv_id":"2305.09617","citing_title":"Towards Expert-Level Medical Question Answering with Large Language Models","ref_index":10,"is_internal_anchor":true},{"citing_arxiv_id":"2402.01411","citing_title":"CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"2403.03952","citing_title":"Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders","ref_index":33,"is_internal_anchor":true},{"citing_arxiv_id":"2407.20595","citing_title":"HALvest-Contrastive: Retrieval-Like Authorship Attribution with Patch-Level Late Interaction","ref_index":8,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7","json":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7.json","graph_json":"https://pith.science/api/pith-number/6FLJ2FRK3565SJWTG3QRIA4TF7/graph.json","events_json":"https://pith.science/api/pith-number/6FLJ2FRK3565SJWTG3QRIA4TF7/events.json","paper":"https://pith.science/paper/6FLJ2FRK"},"agent_actions":{"view_html":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7","download_json":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7.json","view_paper":"https://pith.science/paper/6FLJ2FRK","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2112.11446&json=true","fetch_graph":"https://pith.science/api/pith-number/6FLJ2FRK3565SJWTG3QRIA4TF7/graph.json","fetch_events":"https://pith.science/api/pith-number/6FLJ2FRK3565SJWTG3QRIA4TF7/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7/action/timestamp_anchor","attest_storage":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7/action/storage_attestation","attest_author":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7/action/author_attestation","sign_citation":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7/action/citation_signature","submit_replication":"https://pith.science/pith/6FLJ2FRK3565SJWTG3QRIA4TF7/action/replication_record"}},"created_at":"2026-07-05T03:50:27.025790+00:00","updated_at":"2026-07-05T03:50:27.025790+00:00"}