{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:IEDFAGPWRXT367VCZOSEN4O4FZ","short_pith_number":"pith:IEDFAGPW","schema_version":"1.0","canonical_sha256":"41065019f68de7bf7ea2cba446f1dc2e4e202e8c94af1df944a6cc393ce8a7e4","source":{"kind":"arxiv","id":"2212.13881","version":3},"attestation_state":"computed","paper":{"title":"Mechanism of feature learning in deep fully connected networks and kernel machines that recursively learn features","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","stat.ML"],"primary_cat":"cs.LG","authors_text":"Adityanarayanan Radhakrishnan, Daniel Beaglehole, Mikhail Belkin, Parthe Pandit","submitted_at":"2022-12-28T15:50:58Z","abstract_excerpt":"In recent years neural networks have achieved impressive results on many technological and scientific tasks. Yet, the mechanism through which these models automatically select features, or patterns in data, for prediction remains unclear. Identifying such a mechanism is key to advancing performance and interpretability of neural networks and promoting reliable adoption of these models in scientific applications. In this paper, we identify and characterize the mechanism through which deep fully connected neural networks learn features. We posit the Deep Neural Feature Ansatz, which states that "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2212.13881","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2022-12-28T15:50:58Z","cross_cats_sorted":["cs.AI","stat.ML"],"title_canon_sha256":"de1f89fde0ba64c9e2ae38b7786988bc266af132848b4057a53b1d8ce0267b9e","abstract_canon_sha256":"4ddb39ec6692f89ca797481f33bb9b21e5b39446ca8604aee68f940ea69d8a6c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:08:53.404772Z","signature_b64":"n8O6wuWPy+g+eBiX56wsb79v1w/h50LIgSlNnuopetrwUxnnrnHCCAwI6aTUvE7FVOhLHlyDEOaztxdIXJZjBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"41065019f68de7bf7ea2cba446f1dc2e4e202e8c94af1df944a6cc393ce8a7e4","last_reissued_at":"2026-07-05T06:08:53.404344Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:08:53.404344Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Mechanism of feature learning in deep fully connected networks and kernel machines that recursively learn features","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","stat.ML"],"primary_cat":"cs.LG","authors_text":"Adityanarayanan Radhakrishnan, Daniel Beaglehole, Mikhail Belkin, Parthe Pandit","submitted_at":"2022-12-28T15:50:58Z","abstract_excerpt":"In recent years neural networks have achieved impressive results on many technological and scientific tasks. Yet, the mechanism through which these models automatically select features, or patterns in data, for prediction remains unclear. Identifying such a mechanism is key to advancing performance and interpretability of neural networks and promoting reliable adoption of these models in scientific applications. In this paper, we identify and characterize the mechanism through which deep fully connected neural networks learn features. We posit the Deep Neural Feature Ansatz, which states that "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2212.13881","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2212.13881/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2212.13881","created_at":"2026-07-05T06:08:53.404402+00:00"},{"alias_kind":"arxiv_version","alias_value":"2212.13881v3","created_at":"2026-07-05T06:08:53.404402+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2212.13881","created_at":"2026-07-05T06:08:53.404402+00:00"},{"alias_kind":"pith_short_12","alias_value":"IEDFAGPWRXT3","created_at":"2026-07-05T06:08:53.404402+00:00"},{"alias_kind":"pith_short_16","alias_value":"IEDFAGPWRXT367VC","created_at":"2026-07-05T06:08:53.404402+00:00"},{"alias_kind":"pith_short_8","alias_value":"IEDFAGPW","created_at":"2026-07-05T06:08:53.404402+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.27989","citing_title":"Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17767","citing_title":"Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent","ref_index":148,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15700","citing_title":"AGOP-IxG: A Gradient Covariance Filter for Local Feature Attribution on Tabular Data, with a Controlled Benchmark","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17767","citing_title":"Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent","ref_index":148,"is_internal_anchor":false},{"citing_arxiv_id":"2510.19127","citing_title":"Steering Autoregressive Music Generation with Recursive Feature Machines","ref_index":9,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ","json":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ.json","graph_json":"https://pith.science/api/pith-number/IEDFAGPWRXT367VCZOSEN4O4FZ/graph.json","events_json":"https://pith.science/api/pith-number/IEDFAGPWRXT367VCZOSEN4O4FZ/events.json","paper":"https://pith.science/paper/IEDFAGPW"},"agent_actions":{"view_html":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ","download_json":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ.json","view_paper":"https://pith.science/paper/IEDFAGPW","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2212.13881&json=true","fetch_graph":"https://pith.science/api/pith-number/IEDFAGPWRXT367VCZOSEN4O4FZ/graph.json","fetch_events":"https://pith.science/api/pith-number/IEDFAGPWRXT367VCZOSEN4O4FZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ/action/storage_attestation","attest_author":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ/action/author_attestation","sign_citation":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ/action/citation_signature","submit_replication":"https://pith.science/pith/IEDFAGPWRXT367VCZOSEN4O4FZ/action/replication_record"}},"created_at":"2026-07-05T06:08:53.404402+00:00","updated_at":"2026-07-05T06:08:53.404402+00:00"}