Pith. sign in

Paper Citation Record · LEDGER

Interpreting the Repeated Token Phenomenon in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2503.08908.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.08908 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.044410Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d78de024-8941-40d9-92e2-15d09c2c8a94 · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.949399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:fb6e8412d5a3fc2f6c2f0aa24430f8ba96163a633639206d2131f55affe29ef0

Observation 0e012156-7341-48ac-9b01-fb96607b4c76 · inbound

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability cites this paper.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.044410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.044410Z digest=sha256:3c0263e43ae2fa8c7f67b24e350be80e09f08d23fd124f9c4fb4b512aa77e416

Observation f2700f50-151a-45ab-befb-589ee67e949b · inbound

Taming Outlier Tokens in Diffusion Transformers cites this paper.

Taming Outlier Tokens in Diffusion Transformers Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:07.355579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:20:16.402422Z digest=sha256:e07888f429874dfe8ca63abb23c7a8c2cfb6bb57247cd621054cee2a24313c7d

Observation dd81ed50-db08-4f37-bd51-2a86779a10e6 · inbound

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity cites this paper.

The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:08.552079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:11:04.146711Z digest=sha256:1150e5e006393bf4467663d8d76ee6fa1237106ae9a7720784bb6efef4065b04

Observation fe520b52-5034-4ffe-b4d6-9afde660ccb8 · inbound

Registers Matter for Pixel-Space Diffusion Transformers cites this paper.

Registers Matter for Pixel-Space Diffusion Transformers Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:08:54.081650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:08:23.052621Z digest=sha256:f85f0016e8fb8fc7dc84c0c90220fcd2d13674e7f0b6d7d58dac7afd741d8736

Observation d4833ae4-1148-4208-a113-f82db442c40d · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.712918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:bc646d6b0d6cc38bc9950c44dd375af504b8633321641607eac16f18e442cb85

Observation 7af41a46-775e-400a-8893-099be0bba0aa · inbound

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models cites this paper.

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T10:12:29.133133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:12:29.133133Z digest=sha256:04d05d2fafa73c83b81cfa7b9c30dbbf0638f4dd2672fc8a4c7da7c659f2143e