Pith. sign in

Paper Citation Record · LEDGER

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers

As of 9 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2502.05076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05076 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:25:25.335296Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60c5cb77-9cd5-42ae-a972-8d37de90996d · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.280155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.280155Z digest=sha256:74400c0bbca97c923fc7b40823206f619fe432113713c3bad4078e3a48442ad7

Observation 8807e8bb-2dfb-425a-890a-c20c02444a30 · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.286028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.286028Z digest=sha256:b9d824ab22d356a1a08821f504936ec579df6cf7c13ce4511b443fae356e0de0

Observation 1299dd8a-6a16-473b-90d7-9f8048ac030b · outbound

This paper cites Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.291312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.291312Z digest=sha256:f2750e95e19a574d269e104930a7e95ef1ef3913156a66add87bdbbf260ea95a

Observation 0f9c35b8-c566-4881-a68b-9054ba5c7e74 · outbound

This paper cites A mathematical framework for transformer circuits.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers A mathematical framework for transformer circuits

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.296297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.296297Z digest=sha256:a626464177a8c8208d9a38f61b2e6e5610db67b6ab008d4d9f1efea065cfe00a

Observation a3ef1b7a-3bd0-422e-8172-e43103011287 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Transformer Feed-Forward Layers Are Key-Value Memories

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.301583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.301583Z digest=sha256:d33317860e9e51d4e82c20581989a2db0a9217af1ace00b4e7a11f8cd5be0124

Observation 477fe997-f862-4438-9534-c62270dfa062 · outbound

This paper cites Tensor rank is NP-complete.Journal of algorithms, 11(4):644– 654, 1990.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Tensor rank is NP-complete.Journal of algorithms, 11(4):644– 654, 1990

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:25.528832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:25:25.306914Z digest=sha256:9a0a927969497c93652f7f342b4af8147bb78de5ac039af44f6b01d728f94747

Observation a138a281-cd56-4c8c-9641-a145c6b01daa · outbound

This paper cites Scaling Laws for Fact Memorization of Large Language Models.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Scaling Laws for Fact Memorization of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.312186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.312186Z digest=sha256:b516a5153c4186337965c579dd861c6d7200a681441133aaf5acb4e393ac917e

Observation ada139e0-0b38-48bd-8eff-334963a6f839 · outbound

This paper cites Understanding Factual Recall in Transformers via Associative Memories.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Understanding Factual Recall in Transformers via Associative Memories

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.316854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.316854Z digest=sha256:b8a2b7f7d8556aeab5b3c02c86e6f18958d48bf8cbe5eec8c81c5490834f6a28

Observation 9afedfad-32c4-4bb0-a4c7-289d22a7a2e5 · outbound

This paper cites Tensor rank is hard to approximate.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Tensor rank is hard to approximate

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:25.512580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:25:25.321845Z digest=sha256:511fbcde4fb9190fa3eb211e79402866e643ed43c7651ca35224348a24cea66d

Observation 2f06bb40-e7ff-4fbf-8c07-582594f841d0 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:25.496692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T20:25:25.326351Z digest=sha256:e1b95e4c0bc0f86aab9197dfa65b32551d9b13c83a752257554f1214cb6e8b92

Observation 8f478b2b-1301-4da4-af95-bb85149dbf9f · outbound

This paper cites Breaking the Softmax Bottleneck: A High-Rank RNN Language Model.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.330692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.330692Z digest=sha256:c6f587911019f2890baa20b7c74b241c4ff3ef855ada969580cd5752e581bef6

Observation fd82dd60-88e8-4f6e-9787-82e3c8b5dd33 · outbound

This paper cites Knowledge Circuits in Pretrained Transformers.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Knowledge Circuits in Pretrained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.335296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.335296Z digest=sha256:2474353da0bc96de3accbceaaa15bee503f80e582408796ea5fdc71498486ba3

Pith citing papers

No inbound Pith citation observations are available.