{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:C6HGH2UCGDNPBTJ2OAZH5KJ3FK","short_pith_number":"pith:C6HGH2UC","schema_version":"1.0","canonical_sha256":"178e63ea8230daf0cd3a70327ea93b2a96d29478cbb5d3b686a22345395925d8","source":{"kind":"arxiv","id":"2401.12181","version":1},"attestation_state":"computed","paper":{"title":"Universal Neurons in GPT2 Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Dimitris Bertsimas, Neel Nanda, Qinyi Sun, Tara Rezaei Kheirkhah, Theo Horsley, Wes Gurnee, Will Hathaway, Zifan Carl Guo","submitted_at":"2024-01-22T18:11:01Z","abstract_excerpt":"A basic question within the emerging field of mechanistic interpretability is the degree to which neural networks learn the same underlying mechanisms. In other words, are neural mechanisms universal across different models? In this work, we study the universality of individual neurons across GPT2 models trained from different initial random seeds, motivated by the hypothesis that universal neurons are likely to be interpretable. In particular, we compute pairwise correlations of neuron activations over 100 million tokens for every neuron pair across five different seeds and find that 1-5\\% of"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2401.12181","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-01-22T18:11:01Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"d1c3b8253cd376d29649c1cccd3348f01d06b306714d98ca616443c4b52b75e8","abstract_canon_sha256":"c176443a373aefef669a82ba7d00e90d0df1072c296c32b19ce1fe1ff0ed317e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:36:16.568934Z","signature_b64":"excWR4acoa/a00W8OIKOabgzQDQVcXj05K+UOxa7NQYNNwX0KaU7sZYvSOttXl9DQoh6a0XK1xZsiNSu4+HFAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"178e63ea8230daf0cd3a70327ea93b2a96d29478cbb5d3b686a22345395925d8","last_reissued_at":"2026-07-05T07:36:16.568443Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:36:16.568443Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Universal Neurons in GPT2 Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Dimitris Bertsimas, Neel Nanda, Qinyi Sun, Tara Rezaei Kheirkhah, Theo Horsley, Wes Gurnee, Will Hathaway, Zifan Carl Guo","submitted_at":"2024-01-22T18:11:01Z","abstract_excerpt":"A basic question within the emerging field of mechanistic interpretability is the degree to which neural networks learn the same underlying mechanisms. In other words, are neural mechanisms universal across different models? In this work, we study the universality of individual neurons across GPT2 models trained from different initial random seeds, motivated by the hypothesis that universal neurons are likely to be interpretable. In particular, we compute pairwise correlations of neuron activations over 100 million tokens for every neuron pair across five different seeds and find that 1-5\\% of"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2401.12181","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2401.12181/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2401.12181","created_at":"2026-07-05T07:36:16.568503+00:00"},{"alias_kind":"arxiv_version","alias_value":"2401.12181v1","created_at":"2026-07-05T07:36:16.568503+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2401.12181","created_at":"2026-07-05T07:36:16.568503+00:00"},{"alias_kind":"pith_short_12","alias_value":"C6HGH2UCGDNP","created_at":"2026-07-05T07:36:16.568503+00:00"},{"alias_kind":"pith_short_16","alias_value":"C6HGH2UCGDNPBTJ2","created_at":"2026-07-05T07:36:16.568503+00:00"},{"alias_kind":"pith_short_8","alias_value":"C6HGH2UC","created_at":"2026-07-05T07:36:16.568503+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.02385","citing_title":"How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations","ref_index":62,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12770","citing_title":"WriteSAE: Sparse Autoencoders for Recurrent State","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2408.05147","citing_title":"Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14075","citing_title":"Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12770","citing_title":"WriteSAE: Sparse Autoencoders for Recurrent State","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12770","citing_title":"WriteSAE: Sparse Autoencoders for Recurrent State","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09438","citing_title":"fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2604.04496","citing_title":"The Indra Representation Hypothesis for Multimodal Alignment","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13386","citing_title":"Linear Probe Accuracy Scales with Model Size and Benefits from Multi-Layer Ensembling","ref_index":6,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK","json":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK.json","graph_json":"https://pith.science/api/pith-number/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/graph.json","events_json":"https://pith.science/api/pith-number/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/events.json","paper":"https://pith.science/paper/C6HGH2UC"},"agent_actions":{"view_html":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK","download_json":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK.json","view_paper":"https://pith.science/paper/C6HGH2UC","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2401.12181&json=true","fetch_graph":"https://pith.science/api/pith-number/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/graph.json","fetch_events":"https://pith.science/api/pith-number/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/action/storage_attestation","attest_author":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/action/author_attestation","sign_citation":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/action/citation_signature","submit_replication":"https://pith.science/pith/C6HGH2UCGDNPBTJ2OAZH5KJ3FK/action/replication_record"}},"created_at":"2026-07-05T07:36:16.568503+00:00","updated_at":"2026-07-05T07:36:16.568503+00:00"}