Pith. sign in

Paper Citation Record · LEDGER

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 9 inbound Pith citation observations for arXiv:2507.02559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02559 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.269637Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:11:43.051506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:35:51.305374Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c869a743-df52-4b73-b291-c0d95f896c6b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.056089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.056089Z digest=sha256:276c87353a138159506369f9448dac635b7b4c81fba5a5f82ebb9f82a2ee624f

Observation c0c8e58f-c08b-40df-acbc-810dffbea027 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.172686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.172686Z digest=sha256:f22ae121d370f9efc5e7bba77c466df2a6eb50c444ecefdf55d229ac01c473ea

Observation e041b568-94b6-450a-a69f-aece269aa919 · outbound

This paper cites Why do LLMs attend to the first token?.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Why do LLMs attend to the first token?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.257797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.257797Z digest=sha256:d4f632afba2e4e38b3b5dea5c98a8a8b96eb7109f210cc784e072a1b38628f06

Observation 11fdb6d2-f6eb-4844-bea2-6740f89faba2 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards monosemanticity: Decomposing language models with dictionary learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.344778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.344778Z digest=sha256:5ce0248ed1bad94f5dafa7ffe653d0c415ece30bd288b81e6252211f7ca87698

Observation dde8764f-8c93-4737-b02f-420c9baa97b3 · outbound

This paper cites The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.407102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.407102Z digest=sha256:bf04d6c6a02a8430876891d86490c5ef21f972273cf058727f0e04ddb2e46b2b

Observation 42a3c983-8edd-4529-9f24-b064e7bff7fc · outbound

This paper cites A mathematical framework for transformer circuits.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability A mathematical framework for transformer circuits

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.553041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.553041Z digest=sha256:752cb73a42f8a03467152ec9efb938b964c9ff5b854fd3fa39d3a49291fc169c

Observation e68b7bf3-5b83-4e2e-bd7f-67b328e240c4 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.674321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.674321Z digest=sha256:33715443691965b00b54a7b1603cb1bfbbb708e0540087ab232327ea3f01bdf5

Observation d20ea447-be10-4c78-90b7-ed412a0a938a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.779455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.779455Z digest=sha256:00dd21eda2e2cb55e31026a69ed376697c38d0cbc9d10c2e2cb7a2d06993885e

Observation 798e706f-9568-4c43-9375-68cdc5b1949b · outbound

This paper cites Finding alignments between interpretable causal variables and distributed neural representations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Finding alignments between interpretable causal variables and distributed neural representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.884759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.884759Z digest=sha256:de26fba9ab6a5352eda64eda1c43c0fc874240c1875d1bbec33f818954611c2e

Observation e73c837d-6e25-471b-a3a7-fab05156562f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.963355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.963355Z digest=sha256:28172259d210dc4b01a8aa0dc0151cf0d89feedbcdfd8596504e1d0536bce37d

Observation fb4e697c-81f2-4f85-a760-a354b24aa4d9 · outbound

This paper cites Openwebtext corpus.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.060049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.060049Z digest=sha256:d67c1a241fc12fcb5fa1c714a3b6a0b1338959bd9afb0fb50836a620d74ea915

Observation 111d19bb-32ae-4288-b298-63ef7fc93219 · outbound

This paper cites Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.767788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.177715Z digest=sha256:dba8a338d5c3c9a13b194f81c073fe64814deaa5c86f6d90a0e5295268b17d07

Observation 7239c004-3271-4d79-b952-90e028f687bd · outbound

This paper cites Universal neurons in gpt2 language models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Universal neurons in gpt2 language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.597286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.231561Z digest=sha256:a65a14358a4d703adfb33c11a72ca50eae7df61f5f13a818410c58c2500e80b0

Observation 99434904-f61f-4b34-a984-1539183175a3 · outbound

This paper cites You can remove gpt2's LayerNorm by fine-tuning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability You can remove gpt2's LayerNorm by fine-tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.426689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.288441Z digest=sha256:3bd50f4626196a2a58c39dfb34805c562180ebc887d7e121e29dd43af37c124e

Observation fe9d0fef-a56e-4bdb-be57-c2ff16090900 · outbound

This paper cites How to use and interpret activation patching.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability How to use and interpret activation patching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.371013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.371013Z digest=sha256:97fcceb1608819bacf76e1e9e91d7b14f795003d930204090281a2528a5ae55b

Observation b1b3b626-81be-4d72-90fb-0700bf0324e7 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal covariate shift.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Batch normalization: Accelerating deep network training by reducing internal covariate shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.086538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.500310Z digest=sha256:845bc0ddbef6d227f8fcb1e822ad3de268db34ca5f2da1c534eb612fbca88522

Observation d04f406e-d069-4c5b-9e81-7b1851e40253 · outbound

This paper cites Visit: Visualizing and interpreting the semantic information flow of transformers.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Visit: Visualizing and interpreting the semantic information flow of transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.802792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.581740Z digest=sha256:2d5d2671915f749e097ac8daedb177f58d0f255f16408246f4581eb74c3ad287

Observation a35f9380-5228-4bcf-a980-eda9ca20e0e7 · outbound

This paper cites Sparse autoencoders work on attention layer outputs, Jan 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse autoencoders work on attention layer outputs, Jan 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.597517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.661983Z digest=sha256:9f699df83060cb94a52d31ea312b54d1b674003f90ad50480e5703b9d962b7d1

Observation 2c399d03-4f48-4f46-bda9-872795db513b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.726988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.726988Z digest=sha256:81709a1c3a6f2ddd0f29f5585e01a57b019a35a368dc99280f5961d1a08b3237

Observation 546eccdf-40c9-4f50-8908-d8228a2951a3 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.822219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.822219Z digest=sha256:91156a14aec9c10b17c2d975174002bae15029cba493917b13dc5cd86211131e

Observation d07e6a75-ee99-41cb-86a3-0bb1d54231b9 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Copy Suppression: Comprehensively Understanding an Attention Head

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.921232Z digest=sha256:44dee9d5dcee8738d447cd5f93f2c159100538dcff907f84b855ceb29c5300c5

Observation b3392b0e-1386-4c5d-8410-29acc96e6ab1 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Locating and Editing Factual Associations in GPT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.012912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.012912Z digest=sha256:42972d1c4c548ad0b90bf9c6198edf695e7819ac1598e57b90752805394ede7a

Observation 9be48507-3147-49c3-8d30-9f26f23fa5fe · outbound

This paper cites Tinymodel: A tinystories lm with saes and transcoders, 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Tinymodel: A tinystories lm with saes and transcoders, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.467692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.087598Z digest=sha256:4fe3d7167525ca7ca09fd0ccaa57761969649954e5f7452cd74c01aadae95590

Observation aa89992e-cd0f-4ba6-8fe5-32689f46f113 · outbound

This paper cites Attribution patching: Activation patching at industrial scale, Mar 2023 a.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attribution patching: Activation patching at industrial scale, Mar 2023 a

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.385237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.142216Z digest=sha256:6d0b8269be4008052d77d1a5bae198901be673eb946a43a2251227d077d056a4

Observation 4c84a2a1-d6f1-4065-82e7-f136d1c90939 · outbound

This paper cites Exploratory analysis demo (transformerlens).

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Exploratory analysis demo (transformerlens)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.225301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.226190Z digest=sha256:e87172c823de691853030d5539dbede17fde9c5e15f0941bba73f5e9b74b975e

Observation 1dd69c8d-38aa-4cca-8965-b2d0889424cd · outbound

This paper cites Transformerlens.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformerlens

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.293516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.293516Z digest=sha256:529a64c3085505386f337bd03ea7ba96854bb9f0088c845ae14de120013b9009

Observation f971a940-3747-46ba-bae3-76ae16767369 · outbound

This paper cites interpreting gpt: the logit lens, Aug 2020.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability interpreting gpt: the logit lens, Aug 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.082358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.362337Z digest=sha256:8f5be45a6193fbddf6a8e9ad2da1ae90be9ef0d845f7265230a4401c3fee450d

Observation e749d544-cfd0-4e1b-9617-f80df5827e35 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.418514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.418514Z digest=sha256:5411b06fe41618c1483595ec26e56be6555d267cd4b716b53470cf9390a7e8b9

Observation 4ddc0416-7fcc-431f-8e6a-17f5efd33b43 · outbound

This paper cites Direct and indirect effects.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Direct and indirect effects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.538617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.474033Z digest=sha256:d40ecc1dde2d02b53380240938d01ae1271319712008b9a6529f135a61dbd354

Observation 9a41a79c-d3f2-494b-bef6-e251be332513 · outbound

This paper cites Confidence Regulation Neurons in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Confidence Regulation Neurons in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.531668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.531668Z digest=sha256:cabd6a41d926c12551881c46fe1afd1af2a91cdbe67402f0688c52d2eb7e7217

Observation f1e544b0-9800-4f08-b4d3-7f091708f47a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability LLaMA: Open and Efficient Foundation Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.573752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.573752Z digest=sha256:ef608f88b5402a0c95b2d22753029f40a9a11bfd6c4edc35eb562de9cdffbd0d

Observation 9a14a053-66b1-48d6-9c44-fe69d155270f · outbound

This paper cites Attention is all you need.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attention is all you need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.639147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.639147Z digest=sha256:59335610debd9a80218b3e16a6e0250d322ea1a650714610af1968b48588052d

Observation 2c55ad8d-8a26-4887-b6e2-0b7c0c681830 · outbound

This paper cites Understanding the failure of batch normalization for transformers in nlp.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Understanding the failure of batch normalization for transformers in nlp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.058706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.702651Z digest=sha256:9af0d9448a0fc598c00a53c0eec316127d8ad91f9a231fb67007977f15a78a9c

Observation a1e61622-ebda-4274-882d-000351d125d4 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.798498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.798498Z digest=sha256:8a4f6361e84667a45104a210f272414591f68b800d46454ab43b826530a2cddf

Observation 51304b89-5000-4a5d-90fe-45a9839812db · outbound

This paper cites Re-examining layernorm.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Re-examining layernorm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.869243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.865032Z digest=sha256:0a9071e0b2e3842528241ba5922ea2223edab73fc497b6bdab322bc2ad0016a4

Observation 36a28ddc-1ade-4255-aca7-410308e28a62 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.953864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.953864Z digest=sha256:a1a4c909829266b9fdf9d7c02f7f07f79da10c44fa409a0c256f12ae81893d8b

Observation 0e012156-7341-48ac-9b01-fb96607b4c76 · outbound

This paper cites Interpreting the Repeated Token Phenomenon in Large Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.044410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.044410Z digest=sha256:3c0263e43ae2fa8c7f67b24e350be80e09f08d23fd124f9c4fb4b512aa77e416

Observation 44f04f6d-ef34-414e-be54-aa848e4d1365 · outbound

This paper cites Root mean square layer normalization, 2019.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Root mean square layer normalization, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.138243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.138243Z digest=sha256:4dc072dd06108f53bc23424939931c143ae44aa26496d97cc0ecbc3810861bfd

Observation d1634cc7-c2df-4015-bd8e-792616b6dc9b · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.203317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.203317Z digest=sha256:34a87ea19d645993f2338644f81da316b3a5a3506565f8b7b0326ed90bf683ca

Observation f9d3462e-56da-4ec7-8cf4-b8fc3231803d · outbound

This paper cites Transformers without normalization, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformers without normalization, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.679971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:31:18.269637Z digest=sha256:91537ae1d0d77f6a1d8acad3b2000aa2f43d4de62c062ce177b27e758798af53

Pith citing papers

Observation c1b618d1-fb46-4671-98c7-5601df990771 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:40:50.558061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:f843e48637638efeb4ce3889505a7960308a217732c68fcafe446d889312d0a7

Observation 84b8c5e9-ed7f-4d58-a553-8e19648e3cd2 · inbound

Discovering Interpretable Algorithms by Decompiling Transformers to RASP cites this paper.

Discovering Interpretable Algorithms by Decompiling Transformers to RASP Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:11:43.051506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:11:43.051506Z digest=sha256:75f50f606a5420a610dcbb4dd8205d4f1d1bf8d717309501b57613184b3aca83

Observation 23ef7d84-e166-4f62-8de0-541ef57c2998 · inbound

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers cites this paper.

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.540889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:19:46.400753Z digest=sha256:125e4deb7c6a76bad0d3d2d355ebdad53a4f68489ab51b7f7fd33d7919fd700b

Observation 5ac88c08-06e6-441e-ba65-605d304cf889 · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.362679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:31:17.182481Z digest=sha256:25a09f0ce6228be433e267caded85412eb0ab02df5e7f45b722ec691d16a0d7b

Observation 42cc7a73-a64c-458f-9f99-2e469bff0c8f · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:25.133581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:28:05.507808Z digest=sha256:3363b796255aed478f21b77b590ab57a7eb42cf325ac9dc9303c74c2165de39a

Observation a0fa120d-8d23-4b81-b0b6-15533262c7e2 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:48.064804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:47:13.178911Z digest=sha256:99cf63d7e4732597f3db6ec827a586368e8428492082cac5e384c162b9fa6ef1

Observation b342e9cf-cb1f-4e01-bfaa-8a2cddd8a182 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.932972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.932972Z digest=sha256:fde7d86edaf89b0f518bfc7363d6faf2e46519f04f03aa02f9fd63a236c676dd

Observation a357d3c0-3e46-4594-874e-a9d86d4d0a92 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:35:51.306825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:6df8b7bacdfc13d51d2d958cbbd1531d8466e3913165f8c820ca00e8ff7c47bc

Observation 7dc6d4e6-4bc0-4830-a4e2-274b3eaefe57 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:34:34.461654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:12100109624ca024bd17e14dfba0a3a5fcabb7a9f9bda611fb54e6281ec9be53