{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:FPXUUWGJ7L4JUFGSLMXMAXXF25","short_pith_number":"pith:FPXUUWGJ","schema_version":"1.0","canonical_sha256":"2bef4a58c9faf89a14d25b2ec05ee5d770a2f625b080e9506eea9aa3dfc2d460","source":{"kind":"arxiv","id":"2404.07839","version":2},"attestation_state":"computed","paper":{"title":"RecurrentGemma: Moving Past Transformers for Efficient Open Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Adam Paszke, Alek Andreev, Aleksandar Botev, Andy Brock, Antonia Paterson, Anushan Fernando, Armand Joulin, Arnaud Doucet, Arthur Zucker, Cassidy Hardin, Charlie Chen, Cl\\'ement Farabet, David Budden, David Huntsperger, Demis Hassabis, Elisa Bandy, Evan Senter, George-Cristian Muraru, Glenn Cameron, Guillaume Desjardins, Jenny Brennan, Johan Ferret, Juliette Love, Kathleen Kenealy, Koray Kavukcuoglu, Laurent Sifre, Leonard Berrada, L\\'eonard Hussenot, Ludovic Peran, Luiz Gustavo Martins, Meg Risdal, Mihir Sanjay Kale, Minh Giang, Morgane Rivi\\`ere, Nando de Frietas, Nesh Devanathan, Nilay Chauhan, Noah Fiedel, Olivier Bachem, Paul Mooney, Phil Culliton, Pier Giuseppe Sessa, Pouya Tafti, Raia Hadsell, Raj Gundluru, Razvan Pascanu, Robert Dadashi, Ruba Haroun, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sharad Vikram, Shreya Pathak, Soham De, Srivatsan Srinivasan, Surya Bhupatiraju, Thomas Mesnard, Trevor Gale, Tris Warkentin, Yee Whye Teh, Yutian Chen, Zoubin Ghahramani","submitted_at":"2024-04-11T15:27:22Z","abstract_excerpt":"We introduce RecurrentGemma, a family of open language models which uses Google's novel Griffin architecture. Griffin combines linear recurrences with local attention to achieve excellent performance on language. It has a fixed-sized state, which reduces memory use and enables efficient inference on long sequences. We provide two sizes of models, containing 2B and 9B parameters, and provide pre-trained and instruction tuned variants for both. Our models achieve comparable performance to similarly-sized Gemma baselines despite being trained on fewer tokens."},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2404.07839","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2024-04-11T15:27:22Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"1bdca6fd6fe1dd1db3ced322ce1f5ac9504959361d7e99ab5fd52cad2814b0dd","abstract_canon_sha256":"41c03f0c66337f08ab91edbba0b03a65b1cdea3bff6c4dec67ee83bdcd9b7ea1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:00:03.445136Z","signature_b64":"ZwPPjdgKfldLZmoYF0oZhTRr6t2GlwCuGmsZZ6StO2PIjO+F3U+dkBy6J2yFF+e5H1n9AlbYK5E+Inrj2rT2Dw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2bef4a58c9faf89a14d25b2ec05ee5d770a2f625b080e9506eea9aa3dfc2d460","last_reissued_at":"2026-07-05T09:00:03.444612Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:00:03.444612Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"RecurrentGemma: Moving Past Transformers for Efficient Open Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Adam Paszke, Alek Andreev, Aleksandar Botev, Andy Brock, Antonia Paterson, Anushan Fernando, Armand Joulin, Arnaud Doucet, Arthur Zucker, Cassidy Hardin, Charlie Chen, Cl\\'ement Farabet, David Budden, David Huntsperger, Demis Hassabis, Elisa Bandy, Evan Senter, George-Cristian Muraru, Glenn Cameron, Guillaume Desjardins, Jenny Brennan, Johan Ferret, Juliette Love, Kathleen Kenealy, Koray Kavukcuoglu, Laurent Sifre, Leonard Berrada, L\\'eonard Hussenot, Ludovic Peran, Luiz Gustavo Martins, Meg Risdal, Mihir Sanjay Kale, Minh Giang, Morgane Rivi\\`ere, Nando de Frietas, Nesh Devanathan, Nilay Chauhan, Noah Fiedel, Olivier Bachem, Paul Mooney, Phil Culliton, Pier Giuseppe Sessa, Pouya Tafti, Raia Hadsell, Raj Gundluru, Razvan Pascanu, Robert Dadashi, Ruba Haroun, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sharad Vikram, Shreya Pathak, Soham De, Srivatsan Srinivasan, Surya Bhupatiraju, Thomas Mesnard, Trevor Gale, Tris Warkentin, Yee Whye Teh, Yutian Chen, Zoubin Ghahramani","submitted_at":"2024-04-11T15:27:22Z","abstract_excerpt":"We introduce RecurrentGemma, a family of open language models which uses Google's novel Griffin architecture. Griffin combines linear recurrences with local attention to achieve excellent performance on language. It has a fixed-sized state, which reduces memory use and enables efficient inference on long sequences. We provide two sizes of models, containing 2B and 9B parameters, and provide pre-trained and instruction tuned variants for both. Our models achieve comparable performance to similarly-sized Gemma baselines despite being trained on fewer tokens."},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2404.07839","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2404.07839/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2404.07839","created_at":"2026-07-05T09:00:03.444685+00:00"},{"alias_kind":"arxiv_version","alias_value":"2404.07839v2","created_at":"2026-07-05T09:00:03.444685+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2404.07839","created_at":"2026-07-05T09:00:03.444685+00:00"},{"alias_kind":"pith_short_12","alias_value":"FPXUUWGJ7L4J","created_at":"2026-07-05T09:00:03.444685+00:00"},{"alias_kind":"pith_short_16","alias_value":"FPXUUWGJ7L4JUFGS","created_at":"2026-07-05T09:00:03.444685+00:00"},{"alias_kind":"pith_short_8","alias_value":"FPXUUWGJ","created_at":"2026-07-05T09:00:03.444685+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.02332","citing_title":"Forget Attention: Importance-Aware Attention Is All You Need","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18206","citing_title":"Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2410.13846","citing_title":"LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22416","citing_title":"Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18826","citing_title":"The Routing and Filtering Structure of Attention","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2510.04800","citing_title":"Hybrid Architectures for Language Models: Systematic Analysis and Design Insights","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2501.00663","citing_title":"Titans: Learning to Memorize at Test Time","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03199","citing_title":"Learning the Signature of Memorization in Autoregressive Language Models","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2405.21060","citing_title":"Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality","ref_index":14,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25","json":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25.json","graph_json":"https://pith.science/api/pith-number/FPXUUWGJ7L4JUFGSLMXMAXXF25/graph.json","events_json":"https://pith.science/api/pith-number/FPXUUWGJ7L4JUFGSLMXMAXXF25/events.json","paper":"https://pith.science/paper/FPXUUWGJ"},"agent_actions":{"view_html":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25","download_json":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25.json","view_paper":"https://pith.science/paper/FPXUUWGJ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2404.07839&json=true","fetch_graph":"https://pith.science/api/pith-number/FPXUUWGJ7L4JUFGSLMXMAXXF25/graph.json","fetch_events":"https://pith.science/api/pith-number/FPXUUWGJ7L4JUFGSLMXMAXXF25/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25/action/storage_attestation","attest_author":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25/action/author_attestation","sign_citation":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25/action/citation_signature","submit_replication":"https://pith.science/pith/FPXUUWGJ7L4JUFGSLMXMAXXF25/action/replication_record"}},"created_at":"2026-07-05T09:00:03.444685+00:00","updated_at":"2026-07-05T09:00:03.444685+00:00"}