Pith. sign in

Paper Citation Record · LEDGER

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers

As of 22 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2502.05076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05076 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:25:25.335296Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60c5cb77-9cd5-42ae-a972-8d37de90996d · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.280155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.280155Z digest=sha256:905e66175281c5be13591dcffef5c9c5940bf7dfee84c46fb9ebd733e7c0694c

Observation 8807e8bb-2dfb-425a-890a-c20c02444a30 · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.286028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.286028Z digest=sha256:b9575a570c51c841f0303ff8827dec2cbc985a56caf172521d4f15d5ec62840f

Observation 1299dd8a-6a16-473b-90d7-9f8048ac030b · outbound

This paper cites Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.291312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.291312Z digest=sha256:2e037da5d6ad4d7704d32e800fb47e1558f760d93bf1fe72c3cb501dd20b2322

Observation 0f9c35b8-c566-4881-a68b-9054ba5c7e74 · outbound

This paper cites A mathematical framework for transformer circuits.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers A mathematical framework for transformer circuits

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.296297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.296297Z digest=sha256:adbaee975913361293d6a7bea94bc7e3835d8ee2c78199366822dc02ecb4ae6f

Observation a3ef1b7a-3bd0-422e-8172-e43103011287 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Transformer Feed-Forward Layers Are Key-Value Memories

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.301583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.301583Z digest=sha256:308c1a61f23f30e3615598a78b1f5f8ad8b9c1812165a7d395cfed8c4704814c

Observation 477fe997-f862-4438-9534-c62270dfa062 · outbound

This paper cites Tensor rank is NP-complete.Journal of algorithms, 11(4):644– 654, 1990.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Tensor rank is NP-complete.Journal of algorithms, 11(4):644– 654, 1990

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:25.528832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T20:25:25.306914Z digest=sha256:4f7bc9e9f4e4b8b60d3356568c29ac7f62334c51d12000d54344ef34086b47bf

Observation a138a281-cd56-4c8c-9641-a145c6b01daa · outbound

This paper cites Scaling Laws for Fact Memorization of Large Language Models.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Scaling Laws for Fact Memorization of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.312186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.312186Z digest=sha256:6ae88fa5b161b2880d14c2ae73088582850015e0c4b2ef91702b984eb8202229

Observation ada139e0-0b38-48bd-8eff-334963a6f839 · outbound

This paper cites Understanding Factual Recall in Transformers via Associative Memories.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Understanding Factual Recall in Transformers via Associative Memories

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.316854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.316854Z digest=sha256:ae7352cee8c1728cf989439468162c17852f2d0df7ab2d6ab97a1cb863b18039

Observation 9afedfad-32c4-4bb0-a4c7-289d22a7a2e5 · outbound

This paper cites Tensor rank is hard to approximate.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Tensor rank is hard to approximate

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:25.512580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T20:25:25.321845Z digest=sha256:01da6bd22a4a90a8ffe80bb6f2a81859b6e1164264573c804a70a82d39941e5a

Observation 2f06bb40-e7ff-4fbf-8c07-582594f841d0 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T20:25:25.496692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T20:25:25.326351Z digest=sha256:f743d4ccc0c70ac588e21dca5a9a361a8811473c35b93b7ddd929ff3637faf28

Observation 8f478b2b-1301-4da4-af95-bb85149dbf9f · outbound

This paper cites Breaking the Softmax Bottleneck: A High-Rank RNN Language Model.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.330692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.330692Z digest=sha256:4f59b056bd3a0f6f3a29459a3660d5ac99722f5b18a100e6b4f0f8d166fc257f

Observation fd82dd60-88e8-4f6e-9787-82e3c8b5dd33 · outbound

This paper cites Knowledge Circuits in Pretrained Transformers.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Knowledge Circuits in Pretrained Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.335296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.335296Z digest=sha256:75016c35aa47b15beaeb02aa525eb3a2b83bf2950e9c5e875d4bd061175f891f

Pith citing papers

No inbound Pith citation observations are available.