Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.269637Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 9 inbound Pith citation observations for arXiv:2507.02559.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.269637Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:11:43.051506Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T17:35:51.305374Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c869a743-df52-4b73-b291-c0d95f896c6b · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c8e58f-c08b-40df-acbc-810dffbea027 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e041b568-94b6-450a-a69f-aece269aa919 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Why do LLMs attend to the first token?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11fdb6d2-f6eb-4844-bea2-6740f89faba2 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards monosemanticity: Decomposing language models with dictionary learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde8764f-8c93-4737-b02f-420c9baa97b3 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a3c983-8edd-4529-9f24-b064e7bff7fc · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability A mathematical framework for transformer circuits
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68b7bf3-5b83-4e2e-bd7f-67b328e240c4 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20ea447-be10-4c78-90b7-ed412a0a938a · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798e706f-9568-4c43-9375-68cdc5b1949b · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Finding alignments between interpretable causal variables and distributed neural representations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73c837d-6e25-471b-a3a7-fab05156562f · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Gemini: A Family of Highly Capable Multimodal Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4e697c-81f2-4f85-a760-a354b24aa4d9 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Openwebtext corpus
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111d19bb-32ae-4288-b298-63ef7fc93219 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7239c004-3271-4d79-b952-90e028f687bd · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Universal neurons in gpt2 language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99434904-f61f-4b34-a984-1539183175a3 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability You can remove gpt2's LayerNorm by fine-tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe9d0fef-a56e-4bdb-be57-c2ff16090900 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability How to use and interpret activation patching
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b3b626-81be-4d72-90fb-0700bf0324e7 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Batch normalization: Accelerating deep network training by reducing internal covariate shift
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d04f406e-d069-4c5b-9e81-7b1851e40253 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Visit: Visualizing and interpreting the semantic information flow of transformers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a35f9380-5228-4bcf-a980-eda9ca20e0e7 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse autoencoders work on attention layer outputs, Jan 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c399d03-4f48-4f46-bda9-872795db513b · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 546eccdf-40c9-4f50-8908-d8228a2951a3 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07e6a75-ee99-41cb-86a3-0bb1d54231b9 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Copy Suppression: Comprehensively Understanding an Attention Head
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3392b0e-1386-4c5d-8410-29acc96e6ab1 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Locating and Editing Factual Associations in GPT
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be48507-3147-49c3-8d30-9f26f23fa5fe · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Tinymodel: A tinystories lm with saes and transcoders, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa89992e-cd0f-4ba6-8fe5-32689f46f113 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attribution patching: Activation patching at industrial scale, Mar 2023 a
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c84a2a1-d6f1-4065-82e7-f136d1c90939 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Exploratory analysis demo (transformerlens)
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dd69c8d-38aa-4cca-8965-b2d0889424cd · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformerlens
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f971a940-3747-46ba-bae3-76ae16767369 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability interpreting gpt: the logit lens, Aug 2020
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e749d544-cfd0-4e1b-9617-f80df5827e35 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddc0416-7fcc-431f-8e6a-17f5efd33b43 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Direct and indirect effects
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a41a79c-d3f2-494b-bef6-e251be332513 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Confidence Regulation Neurons in Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e544b0-9800-4f08-b4d3-7f091708f47a · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability LLaMA: Open and Efficient Foundation Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a14a053-66b1-48d6-9c44-fe69d155270f · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attention is all you need
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c55ad8d-8a26-4887-b6e2-0b7c0c681830 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Understanding the failure of batch normalization for transformers in nlp
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1e61622-ebda-4274-882d-000351d125d4 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51304b89-5000-4a5d-90fe-45a9839812db · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Re-examining layernorm
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36a28ddc-1ade-4255-aca7-410308e28a62 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Efficient Streaming Language Models with Attention Sinks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e012156-7341-48ac-9b01-fb96607b4c76 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpreting the Repeated Token Phenomenon in Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f04f6d-ef34-414e-be54-aa848e4d1365 · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Root mean square layer normalization, 2019
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1634cc7-c2df-4015-bd8e-792616b6dc9b · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d3462e-56da-4ec7-8cf4-b8fc3231803d · outbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformers without normalization, 2025
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1b618d1-fb46-4671-98c7-5601df990771 · inbound
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84b8c5e9-ed7f-4d58-a553-8e19648e3cd2 · inbound
Discovering Interpretable Algorithms by Decompiling Transformers to RASP Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ef7d84-e166-4f62-8de0-541ef57c2998 · inbound
Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ac88c08-06e6-441e-ba65-605d304cf889 · inbound
Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42cc7a73-a64c-458f-9f99-2e469bff0c8f · inbound
Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0fa120d-8d23-4b81-b0b6-15533262c7e2 · inbound
Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b342e9cf-cb1f-4e01-bfaa-8a2cddd8a182 · inbound
Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a357d3c0-3e46-4594-874e-a9d86d4d0a92 · inbound
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7dc6d4e6-4bc0-4830-a4e2-274b3eaefe57 · inbound
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.