{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:X5PYK4ILSCEG6CJHQNY6YV2PJH","short_pith_number":"pith:X5PYK4IL","schema_version":"1.0","canonical_sha256":"bf5f85710b90886f09278371ec574f49f5517c2e1ad921293f285beee3e83525","source":{"kind":"arxiv","id":"2203.17189","version":1},"attestation_state":"computed","paper":{"title":"Scaling Up Models and Data with $\\texttt{t5x}$ and $\\texttt{seqio}$","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Aakanksha Chowdhery, Adam Roberts, Afroz Mohiuddin, Aitor Lewkowycz, Alexander Spiridonov, Alexandre Passos, Alex Salcianu, Andrea Gesmundo, Andrew Chen, Anselm Levskaya, Brennan Saeta, Brian Lester, Colin Gaffney, Colin Raffel, Curtis Hawthorne, Dan Garrette, Daniel Andor, Gaurav Mishra, Haitang Hu, Hyung Won Chung, Jacob Austin, James Bradbury, James Lee-Thorp, Jannis Bulian, Jasmijn Bastings, Jeremy Maitin-Shepard, Jianmo Ni, Jonathan H. Clark, Joshua Newlan, Kathleen Kenealy, Livio Baldini Soares, Maarten Bosma, Marc van Zee, Mark Omernick, Marvin Ritter, Noah Fiedel, Noam Shazeer, Ryan Sepassi, Sasha Tsvyashchenko, Sebastian Goodman, Sharan Narang, Stephan Lee, Xavier Garcia","submitted_at":"2022-03-31T17:12:13Z","abstract_excerpt":"Recent neural network-based language models have benefited greatly from scaling up the size of training datasets and the number of parameters in the models themselves. Scaling can be complicated due to various factors including the need to distribute computation on supercomputer clusters (e.g., TPUs), prevent bottlenecks when infeeding data, and ensure reproducible results. In this work, we present two software libraries that ease these issues: $\\texttt{t5x}$ simplifies the process of building and training large language models at scale while maintaining ease of use, and $\\texttt{seqio}$ provi"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2203.17189","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2022-03-31T17:12:13Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"f11d6fdef3d37ff75eaa473b14feb4639b760f44ff276671b16b8d340a5bf1c1","abstract_canon_sha256":"ec771bcec8fae78eafa5871f6bb5e86a2c7513d101d8a9cd52831ebad3e8929c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T04:10:32.353978Z","signature_b64":"UERjute3piKNzLp37VlWAqPwbn+5Ld5av0/f78B7LLklt688GPEp9rjiKS3ZnwJ/yfU3Z3uNiHFVtCU0NpBZBA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"bf5f85710b90886f09278371ec574f49f5517c2e1ad921293f285beee3e83525","last_reissued_at":"2026-07-05T04:10:32.353517Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T04:10:32.353517Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Scaling Up Models and Data with $\\texttt{t5x}$ and $\\texttt{seqio}$","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Aakanksha Chowdhery, Adam Roberts, Afroz Mohiuddin, Aitor Lewkowycz, Alexander Spiridonov, Alexandre Passos, Alex Salcianu, Andrea Gesmundo, Andrew Chen, Anselm Levskaya, Brennan Saeta, Brian Lester, Colin Gaffney, Colin Raffel, Curtis Hawthorne, Dan Garrette, Daniel Andor, Gaurav Mishra, Haitang Hu, Hyung Won Chung, Jacob Austin, James Bradbury, James Lee-Thorp, Jannis Bulian, Jasmijn Bastings, Jeremy Maitin-Shepard, Jianmo Ni, Jonathan H. Clark, Joshua Newlan, Kathleen Kenealy, Livio Baldini Soares, Maarten Bosma, Marc van Zee, Mark Omernick, Marvin Ritter, Noah Fiedel, Noam Shazeer, Ryan Sepassi, Sasha Tsvyashchenko, Sebastian Goodman, Sharan Narang, Stephan Lee, Xavier Garcia","submitted_at":"2022-03-31T17:12:13Z","abstract_excerpt":"Recent neural network-based language models have benefited greatly from scaling up the size of training datasets and the number of parameters in the models themselves. Scaling can be complicated due to various factors including the need to distribute computation on supercomputer clusters (e.g., TPUs), prevent bottlenecks when infeeding data, and ensure reproducible results. In this work, we present two software libraries that ease these issues: $\\texttt{t5x}$ simplifies the process of building and training large language models at scale while maintaining ease of use, and $\\texttt{seqio}$ provi"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2203.17189","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2203.17189/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2203.17189","created_at":"2026-07-05T04:10:32.353577+00:00"},{"alias_kind":"arxiv_version","alias_value":"2203.17189v1","created_at":"2026-07-05T04:10:32.353577+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2203.17189","created_at":"2026-07-05T04:10:32.353577+00:00"},{"alias_kind":"pith_short_12","alias_value":"X5PYK4ILSCEG","created_at":"2026-07-05T04:10:32.353577+00:00"},{"alias_kind":"pith_short_16","alias_value":"X5PYK4ILSCEG6CJH","created_at":"2026-07-05T04:10:32.353577+00:00"},{"alias_kind":"pith_short_8","alias_value":"X5PYK4IL","created_at":"2026-07-05T04:10:32.353577+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.23066","citing_title":"Orbax: Distributed Checkpointing with JAX","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.23066","citing_title":"Orbax: Distributed Checkpointing with JAX","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2204.06745","citing_title":"GPT-NeoX-20B: An Open-Source Autoregressive Language Model","ref_index":79,"is_internal_anchor":false},{"citing_arxiv_id":"2301.13688","citing_title":"The Flan Collection: Designing Data and Methods for Effective Instruction Tuning","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21611","citing_title":"UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15413","citing_title":"Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2211.17192","citing_title":"Fast Inference from Transformers via Speculative Decoding","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2210.02399","citing_title":"Phenaki: Variable Length Video Generation From Open Domain Textual Description","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2311.16867","citing_title":"The Falcon Series of Open Language Models","ref_index":104,"is_internal_anchor":false},{"citing_arxiv_id":"2209.06794","citing_title":"PaLI: A Jointly-Scaled Multilingual Language-Image Model","ref_index":63,"is_internal_anchor":false},{"citing_arxiv_id":"2305.13245","citing_title":"GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2204.02311","citing_title":"PaLM: Scaling Language Modeling with Pathways","ref_index":127,"is_internal_anchor":false},{"citing_arxiv_id":"2403.08295","citing_title":"Gemma: Open Models Based on Gemini Research and Technology","ref_index":95,"is_internal_anchor":false},{"citing_arxiv_id":"2408.00118","citing_title":"Gemma 2: Improving Open Language Models at a Practical Size","ref_index":106,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH","json":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH.json","graph_json":"https://pith.science/api/pith-number/X5PYK4ILSCEG6CJHQNY6YV2PJH/graph.json","events_json":"https://pith.science/api/pith-number/X5PYK4ILSCEG6CJHQNY6YV2PJH/events.json","paper":"https://pith.science/paper/X5PYK4IL"},"agent_actions":{"view_html":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH","download_json":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH.json","view_paper":"https://pith.science/paper/X5PYK4IL","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2203.17189&json=true","fetch_graph":"https://pith.science/api/pith-number/X5PYK4ILSCEG6CJHQNY6YV2PJH/graph.json","fetch_events":"https://pith.science/api/pith-number/X5PYK4ILSCEG6CJHQNY6YV2PJH/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH/action/timestamp_anchor","attest_storage":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH/action/storage_attestation","attest_author":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH/action/author_attestation","sign_citation":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH/action/citation_signature","submit_replication":"https://pith.science/pith/X5PYK4ILSCEG6CJHQNY6YV2PJH/action/replication_record"}},"created_at":"2026-07-05T04:10:32.353577+00:00","updated_at":"2026-07-05T04:10:32.353577+00:00"}