Pith. sign in

Paper Citation Record · LEDGER

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

As of 14 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 9 inbound Pith citation observations for arXiv:2507.02559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02559 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.269637Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:11:43.051506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:35:51.305374Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c869a743-df52-4b73-b291-c0d95f896c6b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.056089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.056089Z digest=sha256:f86c5ffd6d57b132ffc9d4adcd39b2e13ee2e6ad3749a835bc5d45b45b584d06

Observation c0c8e58f-c08b-40df-acbc-810dffbea027 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.172686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.172686Z digest=sha256:8c31c0b979db9538efc813f174304dace03b3c78d2eb502e0bce00b290cf854f

Observation e041b568-94b6-450a-a69f-aece269aa919 · outbound

This paper cites Why do LLMs attend to the first token?.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Why do LLMs attend to the first token?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.257797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.257797Z digest=sha256:0e73cccf3ac2c5f0fa222e75ad6587b124197038558984b3b736604a21a4b7f0

Observation 11fdb6d2-f6eb-4844-bea2-6740f89faba2 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards monosemanticity: Decomposing language models with dictionary learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.344778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.344778Z digest=sha256:efb0f2c557a43fe5cedd45f2e38fea8fdf765b2764ea891c2d456c68f5ed4ba3

Observation dde8764f-8c93-4737-b02f-420c9baa97b3 · outbound

This paper cites The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.407102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.407102Z digest=sha256:92b05dfd14952a4393129bb0a13097dec8977be3085cf3789e33d8f8f6fb0c05

Observation 42a3c983-8edd-4529-9f24-b064e7bff7fc · outbound

This paper cites A mathematical framework for transformer circuits.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability A mathematical framework for transformer circuits

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.553041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.553041Z digest=sha256:770a5784a8dab7faa8e01aa6adf10354375f57d626ebfaab08c1c6f2585441ea

Observation e68b7bf3-5b83-4e2e-bd7f-67b328e240c4 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.674321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.674321Z digest=sha256:4d0e7b262264dd762575d2836e5c9243b0479867be34b2c757d39317ef28e52e

Observation d20ea447-be10-4c78-90b7-ed412a0a938a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.779455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.779455Z digest=sha256:2565beaa4dd4fe36a6d43e844ce290c4b3502dc46dfdbf7779a718221d49735f

Observation 798e706f-9568-4c43-9375-68cdc5b1949b · outbound

This paper cites Finding alignments between interpretable causal variables and distributed neural representations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Finding alignments between interpretable causal variables and distributed neural representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.884759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.884759Z digest=sha256:09cef1b756a73f46fcaeb17d0405aba5bb1fbaf658f30e1aa02645135de47207

Observation e73c837d-6e25-471b-a3a7-fab05156562f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.963355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.963355Z digest=sha256:39a2730b8a83e0f006306dbcefc71eff716fc13ea5a2e2092448dfef87276399

Observation fb4e697c-81f2-4f85-a760-a354b24aa4d9 · outbound

This paper cites Openwebtext corpus.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.060049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.060049Z digest=sha256:2c95f9382f720774401148ddf7983089145e66a0d52496550c0730f87f704930

Observation 111d19bb-32ae-4288-b298-63ef7fc93219 · outbound

This paper cites Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.767788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.177715Z digest=sha256:148b27ffb05e0747bf66d6ece780f4dbe11bc8a08f2a8d150daf572a9e30e5c0

Observation 7239c004-3271-4d79-b952-90e028f687bd · outbound

This paper cites Universal neurons in gpt2 language models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Universal neurons in gpt2 language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.597286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.231561Z digest=sha256:2ae0a0e4c2f6cbcb1afa6f20e2bde62493f2078102a7ae9217ab4992f4850f76

Observation 99434904-f61f-4b34-a984-1539183175a3 · outbound

This paper cites You can remove gpt2's LayerNorm by fine-tuning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability You can remove gpt2's LayerNorm by fine-tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.426689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.288441Z digest=sha256:70ec6717c1098d1a797864e9b11ed29f4e032a29090010beb53361b5c5ca03ae

Observation fe9d0fef-a56e-4bdb-be57-c2ff16090900 · outbound

This paper cites How to use and interpret activation patching.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability How to use and interpret activation patching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.371013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.371013Z digest=sha256:e9d593efcddf6c0dc438cb09a3b5b0b6ac4c91f1139f982c90bf199a229bc1c5

Observation b1b3b626-81be-4d72-90fb-0700bf0324e7 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal covariate shift.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Batch normalization: Accelerating deep network training by reducing internal covariate shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.086538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.500310Z digest=sha256:6ffa3cfa0aea3d1620da896d523f646e0b455a71cfed54d8bb0a0956689aca8f

Observation d04f406e-d069-4c5b-9e81-7b1851e40253 · outbound

This paper cites Visit: Visualizing and interpreting the semantic information flow of transformers.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Visit: Visualizing and interpreting the semantic information flow of transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.802792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.581740Z digest=sha256:3bd8b121401fedd867a427fc2c505e894591e3063244e5bb46b88b61955afb12

Observation a35f9380-5228-4bcf-a980-eda9ca20e0e7 · outbound

This paper cites Sparse autoencoders work on attention layer outputs, Jan 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse autoencoders work on attention layer outputs, Jan 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.597517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.661983Z digest=sha256:42c61e0a45e5adfe8169c73de70bee11c1ea53159825e8754486c62412a2098e

Observation 2c399d03-4f48-4f46-bda9-872795db513b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.726988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.726988Z digest=sha256:6d779e050a67090260253e7a0df00a44dea369e706cec213ad8c61ac9b5d19f9

Observation 546eccdf-40c9-4f50-8908-d8228a2951a3 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.822219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.822219Z digest=sha256:9774274952a8fc48cbd524e46af570298cd2541e9564034598d70fb9f9d0b027

Observation d07e6a75-ee99-41cb-86a3-0bb1d54231b9 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Copy Suppression: Comprehensively Understanding an Attention Head

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.921232Z digest=sha256:c7caa7d3c67e04e089d209a80cea2087149f491089dcd90965d577267e8ad5d1

Observation b3392b0e-1386-4c5d-8410-29acc96e6ab1 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Locating and Editing Factual Associations in GPT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.012912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.012912Z digest=sha256:3eea549fee4e1d0b5101a6feb6694f9767b099a7a835141949af924272d2ea97

Observation 9be48507-3147-49c3-8d30-9f26f23fa5fe · outbound

This paper cites Tinymodel: A tinystories lm with saes and transcoders, 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Tinymodel: A tinystories lm with saes and transcoders, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.467692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.087598Z digest=sha256:364ccc8d73073e49c38904ed10085fe9c580ba835c545e3deddd5762e8003c69

Observation aa89992e-cd0f-4ba6-8fe5-32689f46f113 · outbound

This paper cites Attribution patching: Activation patching at industrial scale, Mar 2023 a.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attribution patching: Activation patching at industrial scale, Mar 2023 a

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.385237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.142216Z digest=sha256:989c9ada97d1f432461d285d6976f761fb9b8cb816b270813492a6e611683923

Observation 4c84a2a1-d6f1-4065-82e7-f136d1c90939 · outbound

This paper cites Exploratory analysis demo (transformerlens).

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Exploratory analysis demo (transformerlens)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.225301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.226190Z digest=sha256:619f588702ef01b3ed406491b3ac09d8dabaa769da6f7b09f6f41238372151b8

Observation 1dd69c8d-38aa-4cca-8965-b2d0889424cd · outbound

This paper cites Transformerlens.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformerlens

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.293516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.293516Z digest=sha256:3e70d62f227ff64416e835e29fa7cd34b153640de9041585afdb87843718dcec

Observation f971a940-3747-46ba-bae3-76ae16767369 · outbound

This paper cites interpreting gpt: the logit lens, Aug 2020.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability interpreting gpt: the logit lens, Aug 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.082358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.362337Z digest=sha256:62866c5612fee24c4d86f9209f620d7fc5191a36af3067f66f26e3607008ecac

Observation e749d544-cfd0-4e1b-9617-f80df5827e35 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.418514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.418514Z digest=sha256:4bb0b997455b105a52da1acf0a928b1531b57a7090ca86d91759b16e6785151a

Observation 4ddc0416-7fcc-431f-8e6a-17f5efd33b43 · outbound

This paper cites Direct and indirect effects.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Direct and indirect effects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.538617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.474033Z digest=sha256:7c9177e889581fa3222947797c17035a7cfbc5afe75ec57d48a50f8d80231509

Observation 9a41a79c-d3f2-494b-bef6-e251be332513 · outbound

This paper cites Confidence Regulation Neurons in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Confidence Regulation Neurons in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.531668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.531668Z digest=sha256:b4a8ab0f3c589c0ba88a0597e6036d4c99ad4a68f713265898b54f2eb9e35182

Observation f1e544b0-9800-4f08-b4d3-7f091708f47a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability LLaMA: Open and Efficient Foundation Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.573752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.573752Z digest=sha256:ce8a40350cef3b0daaeca01754d9228276e97baaf91cf3213df77502889b3441

Observation 9a14a053-66b1-48d6-9c44-fe69d155270f · outbound

This paper cites Attention is all you need.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attention is all you need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.639147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.639147Z digest=sha256:bfba2108c92272e835099c14f347334dcb3ba5f6e24b3a0958f6c8e09685c991

Observation 2c55ad8d-8a26-4887-b6e2-0b7c0c681830 · outbound

This paper cites Understanding the failure of batch normalization for transformers in nlp.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Understanding the failure of batch normalization for transformers in nlp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.058706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.702651Z digest=sha256:e8d22450f8aaa26d968742a59cfa3bce09d31be64a68ac85b67d6684d4967038

Observation a1e61622-ebda-4274-882d-000351d125d4 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.798498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.798498Z digest=sha256:b099fa649cda1e61bfe553641cdd1a34d11ecf292a0846799445f693dec1c352

Observation 51304b89-5000-4a5d-90fe-45a9839812db · outbound

This paper cites Re-examining layernorm.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Re-examining layernorm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.869243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.865032Z digest=sha256:edf71a19a981bd17faa5e8e79abffb15dda8b10fa604a098a1ccea493f3917ca

Observation 36a28ddc-1ade-4255-aca7-410308e28a62 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.953864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.953864Z digest=sha256:7e5face5e34fe6135806a201ba6ab67a5ed918d1e25082ea50a4d14b7d9d02ba

Observation 0e012156-7341-48ac-9b01-fb96607b4c76 · outbound

This paper cites Interpreting the Repeated Token Phenomenon in Large Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.044410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.044410Z digest=sha256:fe5145cf89971c4ed38a47e4830e2500cbd844dcae11c8f4b23c9d885cfaa44c

Observation 44f04f6d-ef34-414e-be54-aa848e4d1365 · outbound

This paper cites Root mean square layer normalization, 2019.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Root mean square layer normalization, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.138243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.138243Z digest=sha256:1b803d4c42fddc6af0eb9305c84f72a1d330a5b4ffa71c4a014ae5d07068fe22

Observation d1634cc7-c2df-4015-bd8e-792616b6dc9b · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.203317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.203317Z digest=sha256:db9a8d4ee9c6788c2061673219fbe2eefe6925bba746092813a16df31682b709

Observation f9d3462e-56da-4ec7-8cf4-b8fc3231803d · outbound

This paper cites Transformers without normalization, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformers without normalization, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.679971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:31:18.269637Z digest=sha256:2a506d73bcf6fcf6f3c737107434241d6ab58b3a9c89340976149243aedec9f6

Pith citing papers

Observation c1b618d1-fb46-4671-98c7-5601df990771 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:40:50.558061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:a67c013efa2657022bdb60b35dd291709551f318c063b1bdd2477ef9488a3cff

Observation 84b8c5e9-ed7f-4d58-a553-8e19648e3cd2 · inbound

Discovering Interpretable Algorithms by Decompiling Transformers to RASP cites this paper.

Discovering Interpretable Algorithms by Decompiling Transformers to RASP Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:11:43.051506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:11:43.051506Z digest=sha256:1807d92d67850209844cc34f5cbff90086f89efddd960fd13c6f89a3b2525de8

Observation 23ef7d84-e166-4f62-8de0-541ef57c2998 · inbound

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers cites this paper.

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.540889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T14:19:46.400753Z digest=sha256:ca4b75a82733dc25d934e2a121454cb6b3182db850c39cff943df1bc27660a42

Observation 5ac88c08-06e6-441e-ba65-605d304cf889 · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.362679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:31:17.182481Z digest=sha256:336478384c1689422de827df20782db589d5075b5302a35dc658de076e8fad33

Observation 42cc7a73-a64c-458f-9f99-2e469bff0c8f · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:25.133581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:28:05.507808Z digest=sha256:8276f437b2fa7d7774650f3189df8b133594e7752dd5dcd1848649c34715101c

Observation a0fa120d-8d23-4b81-b0b6-15533262c7e2 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:48.064804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T17:47:13.178911Z digest=sha256:b22fc4b89f017d2e733f85ec8dc15cd4759d276fb7a480299bc9a86acdd532c0

Observation b342e9cf-cb1f-4e01-bfaa-8a2cddd8a182 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.932972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.932972Z digest=sha256:6d2532cd6480ab1762a5fb5b72d17fef7d1533cd95e56230674da5d037d021fb

Observation a357d3c0-3e46-4594-874e-a9d86d4d0a92 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:35:51.306825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:aa5348ad52f1a511b82750d50119eb93426cbdaa601a3d9a7fd9396625976ff1

Observation 7dc6d4e6-4bc0-4830-a4e2-274b3eaefe57 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:34:34.461654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:801ba67acc264d12ceb4249e0d8b51fdbd082d617bab6e1284bc9ece8e9ac177