{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:2YE4ZXKGPZUHTAA7M2B3DCH3DZ","short_pith_number":"pith:2YE4ZXKG","schema_version":"1.0","canonical_sha256":"d609ccdd467e6879801f6683b188fb1e6dd985781bcdfc3acde9b133e6e3bc4f","source":{"kind":"arxiv","id":"2507.02559","version":1},"attestation_state":"computed","paper":{"title":"Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Galvin Khara, Joachim Schaeffer, Luca Baroni, Marat Subkhankulov, Stefan Heimersheim","submitted_at":"2025-07-03T12:09:04Z","abstract_excerpt":"Layer-wise normalization (LN) is an essential component of virtually all transformer-based large language models. While its effects on training stability are well documented, its role at inference time is poorly understood. Additionally, LN layers hinder mechanistic interpretability by introducing additional nonlinearities and increasing the interconnectedness of individual model components. Here, we show that all LN layers can be removed from every GPT-2 model with only a small increase in validation loss (e.g. +0.03 cross-entropy loss for GPT-2 XL). Thus, LN cannot play a substantial role in"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.02559","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2025-07-03T12:09:04Z","cross_cats_sorted":[],"title_canon_sha256":"b1d197f169531850ca65199f15257dab8595f442fbaa03b06fa88516146c503f","abstract_canon_sha256":"5c5a03a6047b363ff56e3aa92c13d9128bbd449b27d361b095644215be75d66d"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:31:30.864914Z","signature_b64":"yhRaxS6d9Trn45OeaN29X4tKSkEmP7WnBDgvpaDtpDBNsEkX0xnxtHGm8iXUGNAJQUN+lpP9TV/nVkWr0v5gDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"d609ccdd467e6879801f6683b188fb1e6dd985781bcdfc3acde9b133e6e3bc4f","last_reissued_at":"2026-07-05T11:31:30.864554Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:31:30.864554Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Galvin Khara, Joachim Schaeffer, Luca Baroni, Marat Subkhankulov, Stefan Heimersheim","submitted_at":"2025-07-03T12:09:04Z","abstract_excerpt":"Layer-wise normalization (LN) is an essential component of virtually all transformer-based large language models. While its effects on training stability are well documented, its role at inference time is poorly understood. Additionally, LN layers hinder mechanistic interpretability by introducing additional nonlinearities and increasing the interconnectedness of individual model components. Here, we show that all LN layers can be removed from every GPT-2 model with only a small increase in validation loss (e.g. +0.03 cross-entropy loss for GPT-2 XL). Thus, LN cannot play a substantial role in"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.02559","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.02559/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.02559","created_at":"2026-07-05T11:31:30.864611+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.02559v1","created_at":"2026-07-05T11:31:30.864611+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.02559","created_at":"2026-07-05T11:31:30.864611+00:00"},{"alias_kind":"pith_short_12","alias_value":"2YE4ZXKGPZUH","created_at":"2026-07-05T11:31:30.864611+00:00"},{"alias_kind":"pith_short_16","alias_value":"2YE4ZXKGPZUHTAA7","created_at":"2026-07-05T11:31:30.864611+00:00"},{"alias_kind":"pith_short_8","alias_value":"2YE4ZXKG","created_at":"2026-07-05T11:31:30.864611+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.28153","citing_title":"Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models","ref_index":95,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24033","citing_title":"Towards Verifiable Transformers: Solver-Checkable Circuit Explanations","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28153","citing_title":"Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2602.10408","citing_title":"Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2510.23912","citing_title":"Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2604.07098","citing_title":"Selective Neuron Amplification in Transformer Language Models","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2604.07098","citing_title":"Selective Neuron Amplification in Transformer Language Models","ref_index":1,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ","json":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ.json","graph_json":"https://pith.science/api/pith-number/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/graph.json","events_json":"https://pith.science/api/pith-number/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/events.json","paper":"https://pith.science/paper/2YE4ZXKG"},"agent_actions":{"view_html":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ","download_json":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ.json","view_paper":"https://pith.science/paper/2YE4ZXKG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.02559&json=true","fetch_graph":"https://pith.science/api/pith-number/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/graph.json","fetch_events":"https://pith.science/api/pith-number/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/action/storage_attestation","attest_author":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/action/author_attestation","sign_citation":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/action/citation_signature","submit_replication":"https://pith.science/pith/2YE4ZXKGPZUHTAA7M2B3DCH3DZ/action/replication_record"}},"created_at":"2026-07-05T11:31:30.864611+00:00","updated_at":"2026-07-05T11:31:30.864611+00:00"}