{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:2SAIDR7M2WPV7CH7FZDBIZHXK4","short_pith_number":"pith:2SAIDR7M","schema_version":"1.0","canonical_sha256":"d48081c7ecd59f5f88ff2e461464f757251ab620ace18acefaa0fcc821db0179","source":{"kind":"arxiv","id":"2201.02177","version":1},"attestation_state":"computed","paper":{"title":"Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Neural networks can suddenly achieve perfect generalization on small algorithmic tasks long after they have overfitted the training data.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Alethea Power, Harri Edwards, Igor Babuschkin, Vedant Misra, Yuri Burda","submitted_at":"2022-01-06T18:43:37Z","abstract_excerpt":"In this paper we propose to study generalization of neural networks on small algorithmically generated datasets. In this setting, questions about data efficiency, memorization, generalization, and speed of learning can be studied in great detail. In some situations we show that neural networks learn through a process of \"grokking\" a pattern in the data, improving generalization performance from random chance level to perfect generalization, and that this improvement in generalization can happen well past the point of overfitting. We also study generalization as a function of dataset size and f"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":false},"canonical_record":{"source":{"id":"2201.02177","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2022-01-06T18:43:37Z","cross_cats_sorted":[],"title_canon_sha256":"aae4d6517ec56064910e587d7f48ff6067f4834ca66c4d0cf2947700346a86e2","abstract_canon_sha256":"9511d9be986d9d2a086afb5f53cb431ab6087348ade0e71e49051b0cff454e96"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T03:46:30.869032Z","signature_b64":"0uUZRL+L8CN7/dB6L1uC/EHiB2Yz4X68je6WYVun+WvrpGidRWpjUVmSEdYkC9z1KNUswyooVj+vemkHFUkIAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"d48081c7ecd59f5f88ff2e461464f757251ab620ace18acefaa0fcc821db0179","last_reissued_at":"2026-07-05T03:46:30.868533Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T03:46:30.868533Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Neural networks can suddenly achieve perfect generalization on small algorithmic tasks long after they have overfitted the training data.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Alethea Power, Harri Edwards, Igor Babuschkin, Vedant Misra, Yuri Burda","submitted_at":"2022-01-06T18:43:37Z","abstract_excerpt":"In this paper we propose to study generalization of neural networks on small algorithmically generated datasets. In this setting, questions about data efficiency, memorization, generalization, and speed of learning can be studied in great detail. In some situations we show that neural networks learn through a process of \"grokking\" a pattern in the data, improving generalization performance from random chance level to perfect generalization, and that this improvement in generalization can happen well past the point of overfitting. We also study generalization as a function of dataset size and f"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"In some situations we show that neural networks learn through a process of 'grokking' a pattern in the data, improving generalization performance from random chance level to perfect generalization, and that this improvement in generalization can happen well past the point of overfitting.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the grokking behavior observed on these specific small algorithmic datasets reveals a general mechanism of neural network generalization rather than an artifact limited to the chosen tasks, architectures, and optimization regimes.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Neural networks exhibit grokking on small algorithmic datasets, achieving perfect generalization well after overfitting.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Neural networks can suddenly achieve perfect generalization on small algorithmic tasks long after they have overfitted the training data.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"722347474a71b625707310696fdf048a17f2606f7da93d0152416b1e8593c9be"},"source":{"id":"2201.02177","kind":"arxiv","version":1},"verdict":{"id":"722bb388-3eae-4008-9e68-1e1a9619c700","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T19:22:31.706868Z","strongest_claim":"In some situations we show that neural networks learn through a process of 'grokking' a pattern in the data, improving generalization performance from random chance level to perfect generalization, and that this improvement in generalization can happen well past the point of overfitting.","one_line_summary":"Neural networks exhibit grokking on small algorithmic datasets, achieving perfect generalization well after overfitting.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the grokking behavior observed on these specific small algorithmic datasets reveals a general mechanism of neural network generalization rather than an artifact limited to the chosen tasks, architectures, and optimization regimes.","pith_extraction_headline":"Neural networks can suddenly achieve perfect generalization on small algorithmic tasks long after they have overfitted the training data."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2201.02177/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":21,"sample":[{"doi":"","year":null,"title":"Proceedings of the National Academy of Sciences , volume =","work_id":"ec4e299e-f6bc-4097-a042-a9ff15850254","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2006,"title":"Triple descent and the two kinds of overﬁtting: Where & why do they appear? arXiv preprint arXiv:2006.03509","work_id":"a976b8b9-3127-44ae-8c63-5bb1981f52b9","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Universal Transformers","work_id":"8e5baefe-d209-411c-aefc-5acaa9275c8a","ref_index":3,"cited_arxiv_id":"1807.03819","is_internal_anchor":true},{"doi":"","year":null,"title":"Adaptive Computation Time for Recurrent Neural Networks","work_id":"75565443-173e-479c-b0e7-d2464e7630be","ref_index":4,"cited_arxiv_id":"1603.08983","is_internal_anchor":true},{"doi":"","year":null,"title":"Neural Turing Machines","work_id":"0a5ce53c-9670-42b9-8be5-386de7eed50c","ref_index":5,"cited_arxiv_id":"1410.5401","is_internal_anchor":true}],"resolved_work":21,"snapshot_sha256":"aca0d12721867ceeb06e20b696af507876ecb2a27f2b8eceb01e51b33c548201","internal_anchors":7},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2201.02177","created_at":"2026-07-05T03:46:30.868593+00:00"},{"alias_kind":"arxiv_version","alias_value":"2201.02177v1","created_at":"2026-07-05T03:46:30.868593+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2201.02177","created_at":"2026-07-05T03:46:30.868593+00:00"},{"alias_kind":"pith_short_12","alias_value":"2SAIDR7M2WPV","created_at":"2026-07-05T03:46:30.868593+00:00"},{"alias_kind":"pith_short_16","alias_value":"2SAIDR7M2WPV7CH7","created_at":"2026-07-05T03:46:30.868593+00:00"},{"alias_kind":"pith_short_8","alias_value":"2SAIDR7M","created_at":"2026-07-05T03:46:30.868593+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":135,"internal_anchor_count":135,"sample":[{"citing_arxiv_id":"2607.06628","citing_title":"Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06639","citing_title":"At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06648","citing_title":"Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06781","citing_title":"On Explicit Super-Expressive Approximation for Neural Networks","ref_index":76,"is_internal_anchor":true},{"citing_arxiv_id":"2607.08350","citing_title":"Grokking and epoch-wise double descent in quantum neural networks","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2607.08393","citing_title":"Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07066","citing_title":"Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26050","citing_title":"Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining","ref_index":32,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25450","citing_title":"The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25010","citing_title":"Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24396","citing_title":"Parallel Manifold Steering: Efficient Adaptation of Large Associative Memories via Residual Energy Shaping","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25450","citing_title":"The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22873","citing_title":"SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning","ref_index":78,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21228","citing_title":"Sakana Fugu Technical Report","ref_index":117,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21158","citing_title":"Dead-Direction Signatures: A Cheap Spectral Reading of Singular Complexity","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19268","citing_title":"Patnaik-Pearson intrinsic dimension for internal representations of neural networks","ref_index":30,"is_internal_anchor":true},{"citing_arxiv_id":"2606.20737","citing_title":"Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19268","citing_title":"Patnaik-Pearson intrinsic dimension for internal representations of neural networks","ref_index":32,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18164","citing_title":"Learning Dynamics of Chain-of-Thought State Tracking in a Solvable Transformer Model","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01311","citing_title":"From Approximation to Emergence: A Theory of Deep Learning","ref_index":56,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17120","citing_title":"Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12966","citing_title":"Circuit Synchronization Precedes Generalization: A Causal Precursor to Grokking","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11375","citing_title":"When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10877","citing_title":"XtrAIn: Training-Guided Occlusion for Feature Attribution","ref_index":49,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10346","citing_title":"Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning","ref_index":17,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4","json":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4.json","graph_json":"https://pith.science/api/pith-number/2SAIDR7M2WPV7CH7FZDBIZHXK4/graph.json","events_json":"https://pith.science/api/pith-number/2SAIDR7M2WPV7CH7FZDBIZHXK4/events.json","paper":"https://pith.science/paper/2SAIDR7M"},"agent_actions":{"view_html":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4","download_json":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4.json","view_paper":"https://pith.science/paper/2SAIDR7M","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2201.02177&json=true","fetch_graph":"https://pith.science/api/pith-number/2SAIDR7M2WPV7CH7FZDBIZHXK4/graph.json","fetch_events":"https://pith.science/api/pith-number/2SAIDR7M2WPV7CH7FZDBIZHXK4/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4/action/timestamp_anchor","attest_storage":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4/action/storage_attestation","attest_author":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4/action/author_attestation","sign_citation":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4/action/citation_signature","submit_replication":"https://pith.science/pith/2SAIDR7M2WPV7CH7FZDBIZHXK4/action/replication_record"}},"created_at":"2026-07-05T03:46:30.868593+00:00","updated_at":"2026-07-05T03:46:30.868593+00:00"}