Pith. sign in

Paper Citation Record · LEDGER

Transformers without Tears: Improving the Normalization of Self-Attention

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1910.05895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.05895 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:32.086306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T12:34:28.397146Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b62bf687-faa3-4741-ba03-c1d63af740f4 · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Transformers without Tears: Improving the Normalization of Self-Attention

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.535552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:9cf199b114ce968cd484622f490e90ff82acbbcdaee07e7c05af91b25d3c72b8

Observation 22ff5b8e-f7cb-4e49-9e28-ca5b745135f7 · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Transformers without Tears: Improving the Normalization of Self-Attention

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:34:28.399984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:0a17ce1ebfbe8bc5c40ea78bb2eb9c6d591e68f81f7135aa46fe11165942bdd4

Observation e0bc94ef-b68a-4d93-8816-7bd065a5c783 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Transformers without Tears: Improving the Normalization of Self-Attention

Reference 298

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.297137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:e75ef7ce6d601a67d1cec9c6d41a1be87cd26098e75a447ba73b6c0af5c4cffb

Observation c2e63405-5085-4b8d-b3c8-131abe7a4d76 · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Transformers without Tears: Improving the Normalization of Self-Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:32.086306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:32.086306Z digest=sha256:ceb180ca5c671b6dd0f35c5fc1575c774c4284d1cf568864869e23fee72ace3d

Observation f176b608-2526-4295-95c2-1538cb7575ee · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.609473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.609473Z digest=sha256:961deaac2e7da1969981429719fd5039d7cc13b0e8e3e0fd8cc475f9ff91e7e6

Observation c1b805f2-67bb-42a1-9bab-a59829c5c5f5 · inbound

Efficient and Effective Query Context-Aware Learning-to-Rank Model for Sequential Recommendation cites this paper.

Efficient and Effective Query Context-Aware Learning-to-Rank Model for Sequential Recommendation Transformers without Tears: Improving the Normalization of Self-Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:07:00.041601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:07:00.041601Z digest=sha256:c5b2696a0a231cc72259845c7b3039f7990e26a2dba7119ed040cc0a09bb9b16

Observation 0b698c83-5376-4f15-832e-f29ac974cc19 · inbound

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning cites this paper.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Transformers without Tears: Improving the Normalization of Self-Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.074129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.074129Z digest=sha256:4f04e4a448444b3a69e6ec50c191c89260996997e9ac026733e7a98d35246e59

Observation d4414fba-f60a-4526-8073-2b743fece639 · inbound

Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks cites this paper.

Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks Transformers without Tears: Improving the Normalization of Self-Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:20:50.271527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T19:20:50.271527Z digest=sha256:c982e2796aabb34e9048af674c21351e41d31cf38a7ea512279e8f6a3a7244ba

Observation 793975f9-0e15-4db5-b78f-b8d48533e0e6 · inbound

Long-Term Embeddings for Balanced Personalization cites this paper.

Long-Term Embeddings for Balanced Personalization Transformers without Tears: Improving the Normalization of Self-Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.380626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:39:41.045848Z digest=sha256:594680eeab61bedc815a5a7595563a88bada3c63be2a3570a7be9da3cd9c77eb