Pith. sign in

Paper Citation Record · LEDGER

NormFormer: Improved Transformer Pretraining with Extra Normalization

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2110.09456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.09456 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:21:38.284545Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b1d251c3-42a8-4c44-990b-d098de0e89b2 · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:14:25.853489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:cc3cdae7f7dcb4e6e989b32f00e2e5c40ef2dab5ee5d9344c8d3ef558fb35952

Observation 29ff6b07-a487-4db9-b543-7064dd252052 · inbound

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality cites this paper.

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:25.946940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T12:16:25.390683Z digest=sha256:fe46ac6cab2f180e0c2811222602b5742a5100f0f06d1aa8d10536be07766050

Observation b8b59346-a6ea-495f-bd66-e954ea21bd9e · inbound

Learning to (Learn at Test Time): RNNs with Expressive Hidden States cites this paper.

Learning to (Learn at Test Time): RNNs with Expressive Hidden States NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:20:12.251119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T05:20:12.134340Z digest=sha256:71b21d0d1f3d69ebce30c6e784cc943223df28424c87892cad4b3c5e86e5a145

Observation a2fdf57f-fd8a-41ad-bc50-de3f76b65815 · inbound

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers cites this paper.

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:14.296967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:14.296967Z digest=sha256:5088d31628f90322a813f61db1ea0989e73d641e270ebd7c23abecfca7a6b73f

Observation 5b13d27e-365c-4453-a5ce-616fe2a4970d · inbound

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers cites this paper.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.585543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.585543Z digest=sha256:e2189e0d188032d9b67917a2ccf0de3c7ea25bc475255d1dcc86ce733669ce6d

Observation d7398a0b-fb37-4b82-932d-65e3b60f8be7 · inbound

AntLM: Bridging Causal and Masked Language Models cites this paper.

AntLM: Bridging Causal and Masked Language Models NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:40:35.220454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:40:35.220454Z digest=sha256:ce12e441af9273ad911383d08e0de7c94338ea0185c066f8f3ee465f901e9fec

Observation ce1f816d-4e84-412c-ba00-2b1ab475e770 · inbound

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity cites this paper.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.440239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.440239Z digest=sha256:2255e12c8035c95a36e42663f1a9e353b69b27abbc1343962955af587a9dac33

Observation 717c95a1-4c98-4ab5-b7bc-a6cac04c59d2 · inbound

Learning Dynamic Local Context Representations for Infrared Small Target Detection cites this paper.

Learning Dynamic Local Context Representations for Infrared Small Target Detection NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T05:31:31.506044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:31:31.506044Z digest=sha256:6df04cc7f5ca5df48bd12afb1abb450a340cc12c064712902ee3b1f971d0abb2

Observation 192b2961-c928-47f7-861d-87ba0fceeb0a · inbound

Differentially Private Steering for Large Language Model Alignment cites this paper.

Differentially Private Steering for Large Language Model Alignment NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T23:12:56.267577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:12:56.267577Z digest=sha256:2a1639bbe696da3696e75d484e9ace780b4785872bf457a9ee1ab3427872c482

Observation dd02aa21-f87b-4963-999e-5fe241f90d42 · inbound

Spectral-Adaptive Modulation Networks for Visual Perception cites this paper.

Spectral-Adaptive Modulation Networks for Visual Perception NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:35:11.646005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T22:33:48.137960Z digest=sha256:8fd0ffe887a2be6d3c36041710d0f7a862e1f890d01f70d9a1414cd50ed20fd6

Observation e6250802-0980-4c55-8a29-cf60446cfc06 · inbound

Understanding Transformer from the Perspective of Associative Memory cites this paper.

Understanding Transformer from the Perspective of Associative Memory NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.425001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.425001Z digest=sha256:fd53c92b681bbde466508fe098aea543ddc983e5f9706a0453fbfd02c6ab2457

Observation 95eb7219-1915-42ee-8692-65a282ad4ccc · inbound

Residual Matrix Transformers: Scaling the Size of the Residual Stream cites this paper.

Residual Matrix Transformers: Scaling the Size of the Residual Stream NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:41.617138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:10:41.617138Z digest=sha256:912ae27f9d7c9718a316b9373c72d5c01656fe90d18cf5bebc5cb0af6d82d8ad

Observation 0b32a93d-2461-43d1-be7d-49a75c0f1de0 · inbound

Enhancing next token prediction based pre-training for jet foundation models cites this paper.

Enhancing next token prediction based pre-training for jet foundation models NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:13.369717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:13.369717Z digest=sha256:3008f707720ccf828ad9fb9fd8ab7cb64c1448b7f849b06bae8d3383d0876785

Observation 7b20b22d-d874-4795-bebe-a0d7a08359e0 · inbound

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers cites this paper.

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T06:37:53.244800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:37:53.244800Z digest=sha256:7e0982bea27ddeda0ebc51893c13f2bf2958c6f6d0f9c0b3194db16164ca616a

Observation 87d7a6bb-26a6-40c3-8f42-a733a5103058 · inbound

Masked-Token Prediction for Anomaly Detection at the Large Hadron Collider cites this paper.

Masked-Token Prediction for Anomaly Detection at the Large Hadron Collider NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:11:04.347779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T23:24:59.497990Z digest=sha256:cecbd42be9312696a94a8afc26c6aa75c3b60b51d0e55a26d68c6bfcb3ebe32b

Observation 21ab91d6-1470-499f-ab84-275bc0b87d66 · inbound

Dissecting Jet-Tagger Through Mechanistic Interpretability cites this paper.

Dissecting Jet-Tagger Through Mechanistic Interpretability NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:51:22.732285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:49:09.296991Z digest=sha256:1b68284a1c4e48bf71f2d7a4363403cfb85ae294bd82aa22808d475d057a8bae

Observation 269face2-5ba0-4007-9744-f9495e6b9bb8 · inbound

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining cites this paper.

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:27.895465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:12:20.784347Z digest=sha256:cf5d68c1a5f7927826b2dafed0712a273ece2685b07aa63edbe148608f007f03

Observation 41a74ce0-72ff-4cdc-828b-babe7507c3cf · inbound

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models cites this paper.

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:44.043355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T04:27:48.590008Z digest=sha256:e260e4c31dfcaaecae69db63454513f7f1cee11447ee6147c678a958b8cc6664

Observation da89ef92-acd0-43fa-9443-93df1c86e530 · inbound

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry cites this paper.

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-26T05:29:00.032144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T05:22:26.818078Z digest=sha256:92dff9a44145e84cdcd962d770ec7bfc076bc001608ea0794e46996bc0be9269

Observation 900e3749-5352-4242-bdec-812979d4370d · inbound

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth cites this paper.

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:08:35.091316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:08:35.091316Z digest=sha256:b504ae6b8497e403adea153059f705ab7034ec2498025f931f0af0c745b25680

Observation 8af19852-f78d-43d3-a532-5a2890762fd4 · inbound

A Controlled Study of Attention-Only Transformers cites this paper.

A Controlled Study of Attention-Only Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:08.218424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:19:08.218424Z digest=sha256:a8200f5e7b362fe575669462a34c9270bd8e57f71e8d02c7dfc758c5ba0615c7

Observation 4ca5af7c-1300-4e6a-9731-a7264367390f · inbound

Dual Attention Residuals cites this paper.

Dual Attention Residuals NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:36:34.541050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:36:34.541050Z digest=sha256:b345a3b586c4064bfd631f4f6cbc3eee22099efa97cbbfb9001a3d375fd1c086

Observation ceb3623c-9f6b-4a29-9348-c9906ac2fc78 · inbound

Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure cites this paper.

Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:49:13.913766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:49:13.913766Z digest=sha256:da06c7de6667025627f3ad4ef14bb34192390d0cd457dae4c6ae7be0abffeebb

Observation 7a9cb568-1a6c-4a11-92d1-15cde46bcef3 · inbound

Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure cites this paper.

Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T04:21:38.284545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:21:38.284545Z digest=sha256:8717d3d8f1fd8fda1eb433efa1c7260a165a1a6c578b9d7a0d10e658f466e0ac