{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:PW4XVOIWEDRCQUJYZTOQILBMLK","short_pith_number":"pith:PW4XVOIW","schema_version":"1.0","canonical_sha256":"7db97ab91620e2285138ccdd042c2c5a8430083dd4bf47bd34307cb6aaf22109","source":{"kind":"arxiv","id":"2207.08799","version":3},"attestation_state":"computed","paper":{"title":"Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.NE","math.OC","stat.ML"],"primary_cat":"cs.LG","authors_text":"Benjamin L. Edelman, Boaz Barak, Cyril Zhang, Eran Malach, Sham Kakade, Surbhi Goel","submitted_at":"2022-07-18T17:55:05Z","abstract_excerpt":"There is mounting evidence of emergent phenomena in the capabilities of deep learning methods as we scale up datasets, model sizes, and training times. While there are some accounts of how these resources modulate statistical capacity, far less is known about their effect on the computational problem of model training. This work conducts such an exploration through the lens of learning a $k$-sparse parity of $n$ bits, a canonical discrete search problem which is statistically easy but computationally hard. Empirically, we find that a variety of neural networks successfully learn sparse paritie"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2207.08799","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2022-07-18T17:55:05Z","cross_cats_sorted":["cs.NE","math.OC","stat.ML"],"title_canon_sha256":"847ed10f79311e2bb959397bf408d5577edfea427e199f41d1f49c9bc18ea443","abstract_canon_sha256":"f2f0f494a7f3f37b5bc760b1fa54d8757df2e6d2aa0eacad4a2003a369f43aca"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T05:33:14.583785Z","signature_b64":"oJn7LyvKrtua/T6WWVIFifyZO5Bhqu+lTlQwO75gkaP9RglZJ9JOWYXY8kaUmUfra0CNIyQEmFwY1ipI9lhxCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"7db97ab91620e2285138ccdd042c2c5a8430083dd4bf47bd34307cb6aaf22109","last_reissued_at":"2026-07-05T05:33:14.583275Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T05:33:14.583275Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.NE","math.OC","stat.ML"],"primary_cat":"cs.LG","authors_text":"Benjamin L. Edelman, Boaz Barak, Cyril Zhang, Eran Malach, Sham Kakade, Surbhi Goel","submitted_at":"2022-07-18T17:55:05Z","abstract_excerpt":"There is mounting evidence of emergent phenomena in the capabilities of deep learning methods as we scale up datasets, model sizes, and training times. While there are some accounts of how these resources modulate statistical capacity, far less is known about their effect on the computational problem of model training. This work conducts such an exploration through the lens of learning a $k$-sparse parity of $n$ bits, a canonical discrete search problem which is statistically easy but computationally hard. Empirically, we find that a variety of neural networks successfully learn sparse paritie"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2207.08799","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2207.08799/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2207.08799","created_at":"2026-07-05T05:33:14.583338+00:00"},{"alias_kind":"arxiv_version","alias_value":"2207.08799v3","created_at":"2026-07-05T05:33:14.583338+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2207.08799","created_at":"2026-07-05T05:33:14.583338+00:00"},{"alias_kind":"pith_short_12","alias_value":"PW4XVOIWEDRC","created_at":"2026-07-05T05:33:14.583338+00:00"},{"alias_kind":"pith_short_16","alias_value":"PW4XVOIWEDRCQUJY","created_at":"2026-07-05T05:33:14.583338+00:00"},{"alias_kind":"pith_short_8","alias_value":"PW4XVOIW","created_at":"2026-07-05T05:33:14.583338+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.21158","citing_title":"Dead-Direction Signatures: A Cheap Spectral Reading of Singular Complexity","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19542","citing_title":"Tracking Representation Dynamics in Large Language Models with Persistent Homology","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05957","citing_title":"Dead Directions: Geometric Singular Learning","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20314","citing_title":"Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2402.17762","citing_title":"Massive Activations in Large Language Models","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2301.05217","citing_title":"Progress measures for grokking via mechanistic interpretability","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13082","citing_title":"The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2211.00593","citing_title":"Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10237","citing_title":"The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10019","citing_title":"The two clocks and the innovation window: When and how generative models learn rules","ref_index":10,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK","json":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK.json","graph_json":"https://pith.science/api/pith-number/PW4XVOIWEDRCQUJYZTOQILBMLK/graph.json","events_json":"https://pith.science/api/pith-number/PW4XVOIWEDRCQUJYZTOQILBMLK/events.json","paper":"https://pith.science/paper/PW4XVOIW"},"agent_actions":{"view_html":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK","download_json":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK.json","view_paper":"https://pith.science/paper/PW4XVOIW","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2207.08799&json=true","fetch_graph":"https://pith.science/api/pith-number/PW4XVOIWEDRCQUJYZTOQILBMLK/graph.json","fetch_events":"https://pith.science/api/pith-number/PW4XVOIWEDRCQUJYZTOQILBMLK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK/action/storage_attestation","attest_author":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK/action/author_attestation","sign_citation":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK/action/citation_signature","submit_replication":"https://pith.science/pith/PW4XVOIWEDRCQUJYZTOQILBMLK/action/replication_record"}},"created_at":"2026-07-05T05:33:14.583338+00:00","updated_at":"2026-07-05T05:33:14.583338+00:00"}