Pith. sign in

Paper Citation Record · LEDGER

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 9 inbound Pith citation observations for arXiv:2507.02559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02559 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.269637Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:11:43.051506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:35:51.305374Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c869a743-df52-4b73-b291-c0d95f896c6b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.056089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.056089Z digest=sha256:276c87353a138159506369f9448dac635b7b4c81fba5a5f82ebb9f82a2ee624f

Observation c0c8e58f-c08b-40df-acbc-810dffbea027 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.172686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.172686Z digest=sha256:f22ae121d370f9efc5e7bba77c466df2a6eb50c444ecefdf55d229ac01c473ea

Observation e041b568-94b6-450a-a69f-aece269aa919 · outbound

This paper cites Why do LLMs attend to the first token?.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Why do LLMs attend to the first token?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.257797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.257797Z digest=sha256:d4f632afba2e4e38b3b5dea5c98a8a8b96eb7109f210cc784e072a1b38628f06

Observation 11fdb6d2-f6eb-4844-bea2-6740f89faba2 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards monosemanticity: Decomposing language models with dictionary learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.344778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.344778Z digest=sha256:5ce0248ed1bad94f5dafa7ffe653d0c415ece30bd288b81e6252211f7ca87698

Observation dde8764f-8c93-4737-b02f-420c9baa97b3 · outbound

This paper cites The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.407102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.407102Z digest=sha256:bf04d6c6a02a8430876891d86490c5ef21f972273cf058727f0e04ddb2e46b2b

Observation 42a3c983-8edd-4529-9f24-b064e7bff7fc · outbound

This paper cites A mathematical framework for transformer circuits.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability A mathematical framework for transformer circuits

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.553041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.553041Z digest=sha256:752cb73a42f8a03467152ec9efb938b964c9ff5b854fd3fa39d3a49291fc169c

Observation e68b7bf3-5b83-4e2e-bd7f-67b328e240c4 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.674321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.674321Z digest=sha256:33715443691965b00b54a7b1603cb1bfbbb708e0540087ab232327ea3f01bdf5

Observation d20ea447-be10-4c78-90b7-ed412a0a938a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.779455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.779455Z digest=sha256:00dd21eda2e2cb55e31026a69ed376697c38d0cbc9d10c2e2cb7a2d06993885e

Observation 798e706f-9568-4c43-9375-68cdc5b1949b · outbound

This paper cites Finding alignments between interpretable causal variables and distributed neural representations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Finding alignments between interpretable causal variables and distributed neural representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.884759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.884759Z digest=sha256:de26fba9ab6a5352eda64eda1c43c0fc874240c1875d1bbec33f818954611c2e

Observation e73c837d-6e25-471b-a3a7-fab05156562f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.963355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.963355Z digest=sha256:28172259d210dc4b01a8aa0dc0151cf0d89feedbcdfd8596504e1d0536bce37d

Observation fb4e697c-81f2-4f85-a760-a354b24aa4d9 · outbound

This paper cites Openwebtext corpus.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.060049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.060049Z digest=sha256:d67c1a241fc12fcb5fa1c714a3b6a0b1338959bd9afb0fb50836a620d74ea915

Observation 111d19bb-32ae-4288-b298-63ef7fc93219 · outbound

This paper cites Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.767788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.177715Z digest=sha256:ea40fddd43396b1ca6a2dca2dd542853869aba812d8ac69fceafd861ce62be12

Observation 7239c004-3271-4d79-b952-90e028f687bd · outbound

This paper cites Universal neurons in gpt2 language models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Universal neurons in gpt2 language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.597286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.231561Z digest=sha256:d34364f1858219533091f4550e658d783f5c57b7d65736b769175aec99be154d

Observation 99434904-f61f-4b34-a984-1539183175a3 · outbound

This paper cites You can remove gpt2's LayerNorm by fine-tuning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability You can remove gpt2's LayerNorm by fine-tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.426689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.288441Z digest=sha256:a9e01ce9d34f9f880bb133501bbca870cb19738c62906e36c0e633632bb94f7c

Observation fe9d0fef-a56e-4bdb-be57-c2ff16090900 · outbound

This paper cites How to use and interpret activation patching.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability How to use and interpret activation patching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.371013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.371013Z digest=sha256:97fcceb1608819bacf76e1e9e91d7b14f795003d930204090281a2528a5ae55b

Observation b1b3b626-81be-4d72-90fb-0700bf0324e7 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal covariate shift.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Batch normalization: Accelerating deep network training by reducing internal covariate shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.086538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.500310Z digest=sha256:4240431999eda7e0f94b891b4eb9d13d179bd1dec2783d63c6a5d5d05c37aace

Observation d04f406e-d069-4c5b-9e81-7b1851e40253 · outbound

This paper cites Visit: Visualizing and interpreting the semantic information flow of transformers.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Visit: Visualizing and interpreting the semantic information flow of transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.802792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.581740Z digest=sha256:2207da57dbff6022fcba5d2773e49e5d662b1afbc9a28a661be30f1c8a2aaaad

Observation a35f9380-5228-4bcf-a980-eda9ca20e0e7 · outbound

This paper cites Sparse autoencoders work on attention layer outputs, Jan 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse autoencoders work on attention layer outputs, Jan 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.597517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.661983Z digest=sha256:9a9b7eebeab5c4634249d4a865c519b1376d4c88f55888a8e7764a0f5e91d920

Observation 2c399d03-4f48-4f46-bda9-872795db513b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.726988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.726988Z digest=sha256:81709a1c3a6f2ddd0f29f5585e01a57b019a35a368dc99280f5961d1a08b3237

Observation 546eccdf-40c9-4f50-8908-d8228a2951a3 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.822219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.822219Z digest=sha256:91156a14aec9c10b17c2d975174002bae15029cba493917b13dc5cd86211131e

Observation d07e6a75-ee99-41cb-86a3-0bb1d54231b9 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Copy Suppression: Comprehensively Understanding an Attention Head

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.921232Z digest=sha256:44dee9d5dcee8738d447cd5f93f2c159100538dcff907f84b855ceb29c5300c5

Observation b3392b0e-1386-4c5d-8410-29acc96e6ab1 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Locating and Editing Factual Associations in GPT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.012912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.012912Z digest=sha256:42972d1c4c548ad0b90bf9c6198edf695e7819ac1598e57b90752805394ede7a

Observation 9be48507-3147-49c3-8d30-9f26f23fa5fe · outbound

This paper cites Tinymodel: A tinystories lm with saes and transcoders, 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Tinymodel: A tinystories lm with saes and transcoders, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.467692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.087598Z digest=sha256:4d3d383f87abc2744f9fe59f5d84a3a0a5300d1c10491edac6541aaf39df3b75

Observation aa89992e-cd0f-4ba6-8fe5-32689f46f113 · outbound

This paper cites Attribution patching: Activation patching at industrial scale, Mar 2023 a.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attribution patching: Activation patching at industrial scale, Mar 2023 a

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.385237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.142216Z digest=sha256:05a38c4b6f6203cc03d363205b3933ad04f7d254651729f98f360331ed7ade6d

Observation 4c84a2a1-d6f1-4065-82e7-f136d1c90939 · outbound

This paper cites Exploratory analysis demo (transformerlens).

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Exploratory analysis demo (transformerlens)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.225301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.226190Z digest=sha256:888af1e8261f3a3c78bf970ce14ab770ce7a9a46d7f6845b8d62383c248bb7d3

Observation 1dd69c8d-38aa-4cca-8965-b2d0889424cd · outbound

This paper cites Transformerlens.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformerlens

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.293516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.293516Z digest=sha256:529a64c3085505386f337bd03ea7ba96854bb9f0088c845ae14de120013b9009

Observation f971a940-3747-46ba-bae3-76ae16767369 · outbound

This paper cites interpreting gpt: the logit lens, Aug 2020.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability interpreting gpt: the logit lens, Aug 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.082358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.362337Z digest=sha256:0797dfe76b5859c9d8dce4fa230d3aa86b1b70891b0f8d602b763799feab2d91

Observation e749d544-cfd0-4e1b-9617-f80df5827e35 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.418514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.418514Z digest=sha256:5411b06fe41618c1483595ec26e56be6555d267cd4b716b53470cf9390a7e8b9

Observation 4ddc0416-7fcc-431f-8e6a-17f5efd33b43 · outbound

This paper cites Direct and indirect effects.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Direct and indirect effects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.538617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.474033Z digest=sha256:4ef2bcd6e27b2750523b602e3f5128640dbe5419bc7d664b5da073fdbdaa7695

Observation 9a41a79c-d3f2-494b-bef6-e251be332513 · outbound

This paper cites Confidence Regulation Neurons in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Confidence Regulation Neurons in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.531668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.531668Z digest=sha256:cabd6a41d926c12551881c46fe1afd1af2a91cdbe67402f0688c52d2eb7e7217

Observation f1e544b0-9800-4f08-b4d3-7f091708f47a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability LLaMA: Open and Efficient Foundation Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.573752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.573752Z digest=sha256:ef608f88b5402a0c95b2d22753029f40a9a11bfd6c4edc35eb562de9cdffbd0d

Observation 9a14a053-66b1-48d6-9c44-fe69d155270f · outbound

This paper cites Attention is all you need.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attention is all you need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.639147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.639147Z digest=sha256:59335610debd9a80218b3e16a6e0250d322ea1a650714610af1968b48588052d

Observation 2c55ad8d-8a26-4887-b6e2-0b7c0c681830 · outbound

This paper cites Understanding the failure of batch normalization for transformers in nlp.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Understanding the failure of batch normalization for transformers in nlp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.058706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.702651Z digest=sha256:fe3c30b4b21000dfe7230b323e7cd07a0df8abff2e3c6dadb8374fd745007d5a

Observation a1e61622-ebda-4274-882d-000351d125d4 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.798498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.798498Z digest=sha256:8a4f6361e84667a45104a210f272414591f68b800d46454ab43b826530a2cddf

Observation 51304b89-5000-4a5d-90fe-45a9839812db · outbound

This paper cites Re-examining layernorm.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Re-examining layernorm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.869243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.865032Z digest=sha256:020a2741b81c032006b0692d705c4f0272af232648a48c7b5ac706706c4d277e

Observation 36a28ddc-1ade-4255-aca7-410308e28a62 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.953864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.953864Z digest=sha256:a1a4c909829266b9fdf9d7c02f7f07f79da10c44fa409a0c256f12ae81893d8b

Observation 0e012156-7341-48ac-9b01-fb96607b4c76 · outbound

This paper cites Interpreting the Repeated Token Phenomenon in Large Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.044410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.044410Z digest=sha256:3c0263e43ae2fa8c7f67b24e350be80e09f08d23fd124f9c4fb4b512aa77e416

Observation 44f04f6d-ef34-414e-be54-aa848e4d1365 · outbound

This paper cites Root mean square layer normalization, 2019.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Root mean square layer normalization, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.138243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.138243Z digest=sha256:4dc072dd06108f53bc23424939931c143ae44aa26496d97cc0ecbc3810861bfd

Observation d1634cc7-c2df-4015-bd8e-792616b6dc9b · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.203317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.203317Z digest=sha256:8054686d082580b0ba680f6b2231760cecfba948cc6321bbcda37809c14630e5

Observation f9d3462e-56da-4ec7-8cf4-b8fc3231803d · outbound

This paper cites Transformers without normalization, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformers without normalization, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.679971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:31:18.269637Z digest=sha256:a611600a33840bacf7054898fa0c4e947fc244504640fa5dbde3920c03e208a6

Pith citing papers

Observation c1b618d1-fb46-4671-98c7-5601df990771 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:40:50.558061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:65d1bfe57bbc3ce91d96b07fe4d5c0b6fb94c777fb80cee5a9d0c02198904409

Observation 84b8c5e9-ed7f-4d58-a553-8e19648e3cd2 · inbound

Discovering Interpretable Algorithms by Decompiling Transformers to RASP cites this paper.

Discovering Interpretable Algorithms by Decompiling Transformers to RASP Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:11:43.051506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:11:43.051506Z digest=sha256:75f50f606a5420a610dcbb4dd8205d4f1d1bf8d717309501b57613184b3aca83

Observation 23ef7d84-e166-4f62-8de0-541ef57c2998 · inbound

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers cites this paper.

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.540889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:19:46.400753Z digest=sha256:e310bf06223a51e1deb432226435c73d1828099e74fe1ed5bea1dee48bed79f1

Observation 5ac88c08-06e6-441e-ba65-605d304cf889 · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.362679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:31:17.182481Z digest=sha256:aca10537b83d068e8307b4b3866ad8339185c967f900f86da616a32c544c5905

Observation 42cc7a73-a64c-458f-9f99-2e469bff0c8f · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:25.133581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:28:05.507808Z digest=sha256:46740f0b2bb8b508348eae885ebbb4de8bcf1c1d6018777aaadc6a88afe49f0b

Observation a0fa120d-8d23-4b81-b0b6-15533262c7e2 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:48.064804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:47:13.178911Z digest=sha256:28be4425def054cb0e479df44625d4803fc36cdc0f55b23d81209ec4143a366e

Observation b342e9cf-cb1f-4e01-bfaa-8a2cddd8a182 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.932972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.932972Z digest=sha256:fde7d86edaf89b0f518bfc7363d6faf2e46519f04f03aa02f9fd63a236c676dd

Observation a357d3c0-3e46-4594-874e-a9d86d4d0a92 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:35:51.306825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:58ede2eda742e47e4cd14d8af62b82ae64f37cf418115b1c354d16055098a4f6

Observation 7dc6d4e6-4bc0-4830-a4e2-274b3eaefe57 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:34:34.461654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:0f1aad6ac08e37509fd26f6af25861b90aa0621dee8a5860ddd3b896479129d0