{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2026:Y5YZUJBBGHKAOTAPD4EGO5AM43","short_pith_number":"pith:Y5YZUJBB","schema_version":"1.0","canonical_sha256":"c7719a242131d4074c0f1f0867740ce6e61045eca6182b035e1c00390b24b887","source":{"kind":"arxiv","id":"2604.13082","version":2},"attestation_state":"computed","paper":{"title":"The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior","license":"http://creativecommons.org/licenses/by/4.0/","headline":"In encoder-decoder arithmetic models the grokking delay stems from decoder inability to use structure the encoder has already learned.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Laura Gomezjurado Gonzalez","submitted_at":"2026-03-30T22:20:38Z","abstract_excerpt":"Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the source of that delay remains poorly understood. In encoder-decoder arithmetic models, we argue that this delay reflects limited access to already learned structure rather than failure to acquire that structure in the first place. We study one-step Collatz prediction and find that the encoder organizes parity and residue structure within the first few thousand training steps, while output accuracy remains near chance for tens of thousands more. Causa"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":true},"canonical_record":{"source":{"id":"2604.13082","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2026-03-30T22:20:38Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"cfa70a576d760eebe73bf7f49145cf8130d0aba48f6ecd0d92be10be183c2b6b","abstract_canon_sha256":"e5e2c056b38027a2fab43f01d9c02969ae242a2034d6e847b793b3b7ffd5aed9"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-06-19T16:11:23.038551Z","signature_b64":"YkmbmxGzPfW0OwIkHvgMFqbggSz7CZragRfQwdIOhAdwIk5fMRYzt75vjgGqhshh184+4FqJD2I056TGqgQYBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"c7719a242131d4074c0f1f0867740ce6e61045eca6182b035e1c00390b24b887","last_reissued_at":"2026-06-19T16:11:23.038086Z","signature_status":"signed_v1","first_computed_at":"2026-06-19T16:11:23.038086Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior","license":"http://creativecommons.org/licenses/by/4.0/","headline":"In encoder-decoder arithmetic models the grokking delay stems from decoder inability to use structure the encoder has already learned.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Laura Gomezjurado Gonzalez","submitted_at":"2026-03-30T22:20:38Z","abstract_excerpt":"Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the source of that delay remains poorly understood. In encoder-decoder arithmetic models, we argue that this delay reflects limited access to already learned structure rather than failure to acquire that structure in the first place. We study one-step Collatz prediction and find that the encoder organizes parity and residue structure within the first few thousand training steps, while output accuracy remains near chance for tens of thousands more. Causa"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"In encoder-decoder arithmetic models the long delay reflects limited access to already learned structure rather than failure to acquire that structure; transplanting a trained encoder accelerates grokking by 2.75 times while freezing a converged encoder and retraining only the decoder eliminates the plateau and yields 97.6% accuracy.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the causal effects observed in one-step Collatz prediction and the 15-base comparison isolate decoder access as the general cause of grokking delays without confounding factors from architecture, optimization, or task-specific arithmetic.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"The grokking delay in encoder-decoder models on one-step Collatz prediction stems from decoder inability to use early-learned encoder representations of parity and residue structure, with numeral base acting as a strong inductive bias that can raise accuracy from failure to 99.8%.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"In encoder-decoder arithmetic models the grokking delay stems from decoder inability to use structure the encoder has already learned.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"c2333d7385f3cd64535701213887144874ff2c5515f17bb6639b74e86198d01f"},"source":{"id":"2604.13082","kind":"arxiv","version":2},"verdict":{"id":"3cba6d2b-9b36-4de6-b9d3-2720973992d8","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-14T21:12:43.869799Z","strongest_claim":"In encoder-decoder arithmetic models the long delay reflects limited access to already learned structure rather than failure to acquire that structure; transplanting a trained encoder accelerates grokking by 2.75 times while freezing a converged encoder and retraining only the decoder eliminates the plateau and yields 97.6% accuracy.","one_line_summary":"The grokking delay in encoder-decoder models on one-step Collatz prediction stems from decoder inability to use early-learned encoder representations of parity and residue structure, with numeral base acting as a strong inductive bias that can raise accuracy from failure to 99.8%.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the causal effects observed in one-step Collatz prediction and the 15-base comparison isolate decoder access as the general cause of grokking delays without confounding factors from architecture, optimization, or task-specific arithmetic.","pith_extraction_headline":"In encoder-decoder arithmetic models the grokking delay stems from decoder inability to use structure the encoder has already learned."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.13082/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"2c30a2e679ae0e15d1230543e7b30d0ff5e4069babb0a7a6cbf65dd496aaa51c"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2604.13082","created_at":"2026-06-19T16:11:23.038146+00:00"},{"alias_kind":"arxiv_version","alias_value":"2604.13082v2","created_at":"2026-06-19T16:11:23.038146+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2604.13082","created_at":"2026-06-19T16:11:23.038146+00:00"},{"alias_kind":"pith_short_12","alias_value":"Y5YZUJBBGHKA","created_at":"2026-06-19T16:11:23.038146+00:00"},{"alias_kind":"pith_short_16","alias_value":"Y5YZUJBBGHKAOTAP","created_at":"2026-06-19T16:11:23.038146+00:00"},{"alias_kind":"pith_short_8","alias_value":"Y5YZUJBB","created_at":"2026-06-19T16:11:23.038146+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2605.20441","citing_title":"Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics","ref_index":12,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43","json":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43.json","graph_json":"https://pith.science/api/pith-number/Y5YZUJBBGHKAOTAPD4EGO5AM43/graph.json","events_json":"https://pith.science/api/pith-number/Y5YZUJBBGHKAOTAPD4EGO5AM43/events.json","paper":"https://pith.science/paper/Y5YZUJBB"},"agent_actions":{"view_html":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43","download_json":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43.json","view_paper":"https://pith.science/paper/Y5YZUJBB","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2604.13082&json=true","fetch_graph":"https://pith.science/api/pith-number/Y5YZUJBBGHKAOTAPD4EGO5AM43/graph.json","fetch_events":"https://pith.science/api/pith-number/Y5YZUJBBGHKAOTAPD4EGO5AM43/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43/action/timestamp_anchor","attest_storage":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43/action/storage_attestation","attest_author":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43/action/author_attestation","sign_citation":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43/action/citation_signature","submit_replication":"https://pith.science/pith/Y5YZUJBBGHKAOTAPD4EGO5AM43/action/replication_record"}},"created_at":"2026-06-19T16:11:23.038146+00:00","updated_at":"2026-06-19T16:11:23.038146+00:00"}