Pith. sign in

Paper Citation Record · LEDGER

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

As of 17 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2608.07921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07921 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:46:12.923935Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9571e9c8-936b-4993-a546-b02d2972f964 · outbound

This paper cites Attention is all you need,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.818591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.818591Z digest=sha256:d4baa3a2c9fa375772ef3f645a55b35cd5e02100427c36d75cfaba13f98e504b

Observation ffaa6dd1-86fc-4e29-ab75-879b7e201995 · outbound

This paper cites A mathematical framework for transformer circuits,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention A mathematical framework for transformer circuits,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.376120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.823516Z digest=sha256:144c3cb13ae10dd041192ffef801decb7863140ff3253597b0c5c7fdad92530c

Observation 6df36632-a5b4-4726-b453-4d37bfc58897 · outbound

This paper cites In-context Learning and Induction Heads.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention In-context Learning and Induction Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.828334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.828334Z digest=sha256:7e27c52467e67c00cb728a1f1a49c49c90bf5f1ef0c02d651e9bd7c5cbbfc857

Observation 1fc4c0f1-1af0-41bc-9330-f0e1bb4c0aa2 · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention LoRA: Low-rank adaptation of large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.833310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.833310Z digest=sha256:0fbee9f48e6db739d7eed24c5c7c97b0957e82d2a40e62a083ac4193c424ffe6

Observation 51ae4dc3-e53f-493e-b8cb-487804b38312 · outbound

This paper cites Distribution of eigenvalues for some sets of random matrices,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Distribution of eigenvalues for some sets of random matrices,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.837985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.837985Z digest=sha256:182b1a76468f5234d08d5ae4cebbb8a981798dd3190db228c7ffaefe575e4f6e

Observation bfa875f8-b8f7-456a-aff7-d05b9934f043 · outbound

This paper cites Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.842982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.842982Z digest=sha256:0156f632d211da3d87a9132f32332cdeb566303be656aad6a87d0f517d2b4ac3

Observation aa812edd-7087-4962-a5eb-600c48a76158 · outbound

This paper cites A Spectral Condition for Feature Learning.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention A Spectral Condition for Feature Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.847768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.847768Z digest=sha256:601963b6099b48d1d644e513ca24f837320790ed7185acdc38679533bbdef37f

Observation 1aa839ff-5760-4f56-a6eb-fcd09ed7ae0d · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.340844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.852294Z digest=sha256:8ab44a06ecb93dfd20576ce3d255830d60649f3de112141885e4b31284b4aac6

Observation 5271e91f-6d3d-4c87-af8d-f5716557cfc1 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.856561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.856561Z digest=sha256:854d85860517f055a7b2484cfb888b674966fb5c7e668d1fd3fb8fa7bf980ee6

Observation d90c9779-4b3a-420e-a4e3-be695088610d · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention OPT: Open Pre-trained Transformer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.861338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.861338Z digest=sha256:b7c06a4614e941de1b2947a26e1d5607ec794a9e491a6466ff9078e4563ffe82

Observation b706274c-7db4-411a-b38f-b1f4a7f05a81 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention LLaMA: Open and Efficient Foundation Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.866008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.866008Z digest=sha256:4abccc13f98977d9fb6a47aa67097a336267c004e8928172966e751ab1d39c5a

Observation 2413c58a-76d3-405a-8d68-922c590bb004 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.870550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.870550Z digest=sha256:e4a17d98306fd78c86e6c40668c78ea7677879cad71616b664b3ae4accbc05e3

Observation acbdc865-ee55-4dff-9ef1-746a74c985eb · outbound

This paper cites The Llama 3 Herd of Models.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.875200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.875200Z digest=sha256:7fa684b4611488ba6a0d300bd2301dce4ecc8df45fc070f82ee1245125dfa55e

Observation 12a788cf-30a9-41a0-9445-0525ed4be4ff · outbound

This paper cites Mistral 7B.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.880012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.880012Z digest=sha256:b6bacaa07fc2a80b638e52b75bda53d7793854120cbc66ff4cc2741c4859eef0

Observation b9d15885-29bd-4e58-8c34-21e151a7e8e8 · outbound

This paper cites Qwen2.5 Technical Report.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.884531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.884531Z digest=sha256:f2298a29451112074ef6bc142464bc93334e8d1f1e206b24b0a895eac7613b8f

Observation 0576c148-1bb6-462b-904f-d5484ec6e6a5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.889024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.889024Z digest=sha256:1d08e8202eb3afbb76e198c4b7713de7f32c0622b6ec0c2e91a9b3c3f244ab96

Observation 709b51fb-a74c-4622-87fd-795a82554477 · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention GQA: Training generalized multi-query transformer models from multi-head checkpoints,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.324435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.893597Z digest=sha256:5b9530d964549a57e5cd2885c0abce031b1f6ead15d6983a434846fe65ed15f3

Observation bb151c8b-b6f2-46c8-8227-f247e94e8f7c · outbound

This paper cites HellaSwag: Can a machine really finish your sentence?,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention HellaSwag: Can a machine really finish your sentence?,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.308951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.897969Z digest=sha256:4c606f4928c283ab290ed5553af9ff18d1a9e6d8f70c4960142b51f915d0a7c5

Observation 7b8fb840-ab0f-45c6-a409-f5c229fc15c2 · outbound

This paper cites Measuring massive multitask language understanding,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Measuring massive multitask language understanding,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.294633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.902307Z digest=sha256:82013e0a74d91f86b5560e5b6ff4cd552e1f1691f2be9d2ea26ac067c4219ed2

Observation 677e8a3c-4f5b-49ad-861d-28f79bb569bc · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention PIQA: Reasoning about physical commonsense in natural language,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.279462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.906552Z digest=sha256:90c17e6aae80194ae6ce4e2bcacbc6184c57f75f0be8197fa6f3620bba1e04a3

Observation 59a8edf4-fac9-45ab-bd2a-b2b9f2a985e5 · outbound

This paper cites A framework for few-shot language model evaluation,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention A framework for few-shot language model evaluation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.910787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.910787Z digest=sha256:51ad1eadf41b0a90eacdf20f27428c6d8be7ad10b0fe7620d9e768ea50e3cea8

Observation 5791b5aa-6aaa-4ccc-b43e-cb3b0c46ef0f · outbound

This paper cites What does BERT learn about the structure of language?,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention What does BERT learn about the structure of language?,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:46:13.264492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.915395Z digest=sha256:cb9ebc872ab445af53806712d93c1ea6f63ff3a970f60eb1113c942974f5e4e4

Observation 3b51e1c5-c4c5-4c27-b2a0-d39643dd88ee · outbound

This paper cites Toy Models of Superposition.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention Toy Models of Superposition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:46:12.919561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:46:12.919561Z digest=sha256:5118922dbeb9d724ca0b26175464a0560e9b9e1dd8d23648ca8010c5b0d4c0a5

Observation 0d513932-740e-473f-8f61-a70b76248cae · outbound

This paper cites LORA-CRAFT: Cross-layer rank adaptation via frozen Tucker decomposition of pre- trained attention weights,.

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention LORA-CRAFT: Cross-layer rank adaptation via frozen Tucker decomposition of pre- trained attention weights,

Reference 24

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:46:13.094666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:46:12.923935Z digest=sha256:e4d0fddba31eb6881803f7d7f7d77e78385b8c4f3a17a2ad09e9eb868d05120f

Pith citing papers

No inbound Pith citation observations are available.