Pith. sign in

Paper Citation Record · LEDGER

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 11 inbound Pith citation observations for arXiv:2506.01115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01115 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:22.908318Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.096284Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.880747Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8363ba87-329c-42c9-9ecc-0a23659bea68 · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.571092Z digest=sha256:0d08c8a3cbd4050600e98f08e34c52f471b591c7509c9fcb95aa735ae65b4a77

Observation e9abca2d-5c2b-440f-9692-f09073b1fa79 · outbound

This paper cites The Curious Case of Benign Memorization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The Curious Case of Benign Memorization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.702593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.577524Z digest=sha256:2b38f4e0223c9c7dfd4f5706c6641a1fcc9ad98c02492654ef08090f155ed047

Observation 73175888-b0e0-4f1b-a063-df84c3ac5ce3 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.158136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.583250Z digest=sha256:023851bfc14967642c792212acf9a7d46918c41fb0ea30e0584b750390c6c430

Observation e41ea57a-c370-4585-bf0c-5511b8b9b17a · outbound

This paper cites A closer look at memorization in deep networks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A closer look at memorization in deep networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.143410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.588325Z digest=sha256:c7c59323c67af03dec744c069603adc797b63875a02394e38aa835a2426fbd84

Observation 117151e8-5cca-4bda-b64b-6e54be0079b1 · outbound

This paper cites Scaling mlps: A tale of inductive bias.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Scaling mlps: A tale of inductive bias

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.123851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.593277Z digest=sha256:925142bf18b93274fcc3167271558db743c8c7e62870f0b028b757bc9e91fbc2

Observation 879bb15d-bee3-410b-a93f-0189690761f1 · outbound

This paper cites Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.599011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.599011Z digest=sha256:530ac6797ba46d401669890965d32706178a57c52e08c4ee19b1e4018d4d354c

Observation 0ea2cbf4-4bb5-4b85-af8e-f660e06859a4 · outbound

This paper cites Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.096487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.605192Z digest=sha256:867d9006f5c0c3ebcf105c5bbbeceec0b454b6390726d6f201c2e5e0a897d4ea

Observation 662cb06e-280c-496c-b623-4dc346d22452 · outbound

This paper cites Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.631525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.611391Z digest=sha256:0f17bacc913560fc660ba6adcee7728b54408b9edc36adf2cb28fe1a8f8db539

Observation 3191c054-6843-42be-a1a1-84be8dbd1f98 · outbound

This paper cites Transformers generalize differently from information stored in context vs in weights.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers generalize differently from information stored in context vs in weights

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.619625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.619625Z digest=sha256:5f42ac87e63bf1856781b1ca3c7063c67da8fe88bbd4bce5ff1bc95cd8ba15d9

Observation 30c7a563-58a0-4209-ae7e-594a24231996 · outbound

This paper cites Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.082310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.626924Z digest=sha256:1c4511dae928aa72e5e3eb0bb9955224c24b6542c2b771bd8b3701fc8e421a93

Observation 68fa1850-9619-4b11-837e-0cf11690ca00 · outbound

This paper cites Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.068319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.633098Z digest=sha256:59da794ea3737ffc660aeb11b1b1200ab1beefce38faa1fb68f44fef351eeb8d

Observation 8ca9d54d-76bd-4019-af87-fba6bb10bc3e · outbound

This paper cites Summing up the facts: Additive mechanisms behind factual recall in llms, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Summing up the facts: Additive mechanisms behind factual recall in llms, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.054015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.639502Z digest=sha256:823b676ff6b9cb83b0ad151e4e6f347cdc568c71e1226585717792064b853b41

Observation 1e846ff2-6ed8-4deb-8034-b1fb20869d90 · outbound

This paper cites Induction heads as an essential mechanism for pattern matching in in-context learning.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Induction heads as an essential mechanism for pattern matching in in-context learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.040339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.647499Z digest=sha256:d81a8acae46dfa4df85f34a9917e0a49d09d09e9febb29dc14a3ef03f5c1d222

Observation 83408a97-f0ab-447b-a5e6-c01f0c6a15c1 · outbound

This paper cites Knowledge neurons in pretrained transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge neurons in pretrained transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.655162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.655162Z digest=sha256:1e1f2a169eb3daea94c422a21a1cffd0c0d04d6978ebedff79daa19cd7eac8e7

Observation 44cc6a85-ddda-418d-9e36-2645fb4adb39 · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.027077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.661558Z digest=sha256:7bd30879e46334d044c35e928b5c26039a99cb552becc381a538068b923a2e9d

Observation 28cc1dae-114a-43c7-a596-2b6d94072207 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.667162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.667162Z digest=sha256:8a484ca0b050cb885fd4a0497b5083394ea02a37ad5d2a7548b380ffe3dceef1

Observation 1d511882-7077-4c50-b2ab-5aa70f144bfa · outbound

This paper cites Edelman, eran malach, and Surbhi Goel.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Edelman, eran malach, and Surbhi Goel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.008710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.672227Z digest=sha256:01b6ace4c8d60fd1c6f85c7c716559a63e07b690dc1d1849444ab404bbca469c

Observation 6de929e4-298d-4f9f-bd46-dcf17694bbd6 · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.989975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.677623Z digest=sha256:1327a4f0f5cbdfb1990f91372f54e86bda7629fa4d8ceeeaac409f71e1b213cf

Observation 58bf9a70-406d-4f25-8238-7eedfc9358da · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer Feed-Forward Layers Are Key-Value Memories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.682505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.682505Z digest=sha256:40df797daeb45b742e4d4b7186a082bfedf874c1bf831394c76577f2b48fd435

Observation c47962c0-f560-41db-b6ee-fa5f74c6d609 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers are key-value memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.687732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.687732Z digest=sha256:829211ee9e3f4e8b92e2201efdf53cfb03889f2cb63b1e1bf34471a00ed56064

Observation c0f87075-542c-4a28-afc8-718b2ecba1d3 · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.692163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.692163Z digest=sha256:a8f3a505431642bde5ad3b8b3c2774d05f2640c209a6c4449a437053e5bdf7b9

Observation 2dc95dbb-4299-4414-a852-483b5d696015 · outbound

This paper cites Dissectingrecalloffactualassociations inauto-regressivelanguagemodels.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Dissectingrecalloffactualassociations inauto-regressivelanguagemodels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.696572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.696572Z digest=sha256:7e4ff909a3865f9d9d7a907c392e5bc9de8894c5991ad261633c76a8981a6fb7

Observation a3e0fa06-aee0-48fa-b920-79811df9a4f2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.701367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.701367Z digest=sha256:3779b4886a46546d9688f320dfe11c3e882156e33821164721daa586b9c0070f

Observation 9631f467-18cd-46a7-aa33-fb364dfc04cf · outbound

This paper cites Smith, and Roy Schwartz.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Smith, and Roy Schwartz

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.974388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.705870Z digest=sha256:b26d798dfc2d9d2461f4bc025c575f568f5419bae754bda6e54f5432c9f74873

Observation 2849f7ac-531f-4d2e-9980-80ffde91facd · outbound

This paper cites Simplifying Transformer Blocks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Simplifying Transformer Blocks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.710780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.710780Z digest=sha256:34fed603f01df2e0c112ef3abc6d606cab8258b7eb4cd00987958d7630ecb8b9

Observation 20848180-274a-4d69-a981-68c6b2675f35 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.716704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.716704Z digest=sha256:a4b216732ca58129ca81cef0db62a1fea31a40f36396f487d06993b3c1165f58

Observation 547444a1-677d-4f56-a61f-815b88a3eef3 · outbound

This paper cites What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.722380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.722380Z digest=sha256:f65148dee4c27f9d97207902b5a07ee5b871c78844fb6170f3f166c16116f427

Observation 657d772e-82b4-4721-82b6-807cdc2c3692 · outbound

This paper cites Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.960220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.728573Z digest=sha256:804c39e9387e8a8636827dbc8e3d07443182f70aa888c42cd57f7e838e6bc2ee

Observation 1487f04a-6d72-4d38-87fd-03cc49bd03a7 · outbound

This paper cites Deep Neural Networks as Gaussian Processes.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Neural Networks as Gaussian Processes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.735092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.735092Z digest=sha256:e9ee8e58b6c843adf55f8d647ad055be957bc140b8131ec98e30d7327468ca9d

Observation 712c06fc-030a-4711-a0f1-dbb80d30ae53 · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer FNet: Mixing Tokens with Fourier Transforms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.740900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.740900Z digest=sha256:2fb608a6e7504bdfc280c1a8cdf4854047bea7812f4d7bef7d5b9350b63c1715

Observation 5c2b73b7-af45-463d-9a16-cae3f6618d03 · outbound

This paper cites The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.946050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.745978Z digest=sha256:31e54909815b866619179b4f448e79aa1dccb511954fb867b8f0f30efdb781fd

Observation 59b06326-5916-47b7-9a17-d49c8b382d79 · outbound

This paper cites Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.750524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.750524Z digest=sha256:fb4083712c5bfdcccac4523aa845f7c37054eb3091faaf5409e9c40d717daf56

Observation f8d18586-9ccc-4f75-92cb-096e3db23870 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating and Editing Factual Associations in GPT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.755732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.755732Z digest=sha256:d954ba129a4bfe6f720702fa63c13b6fedbc0ca0903f839ca361368c13b37da8

Observation a0aca01b-a1f7-4fd7-a5da-eabfdb3d31f5 · outbound

This paper cites Pointer Sentinel Mixture Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Pointer Sentinel Mixture Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.760917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.760917Z digest=sha256:77bdc2a9ad8285493b0e5f8d9062f2c85350b4d1cfbcc655ea8af2cc90c4c0d6

Observation 1565742c-8d44-47e6-a5d7-46ce1da0bc03 · outbound

This paper cites Language models implement simple word2vec-style vector arithmetic, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Language models implement simple word2vec-style vector arithmetic, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.927859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.765913Z digest=sha256:1c0787f0ddc672b6ae6ff423b103a67711cd3499d4bf05f20fff57f307afc652

Observation 80ff018a-334f-4a56-abb3-99e13bcdbd6a · outbound

This paper cites Universal approximation property of random neural networks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Universal approximation property of random neural networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.911077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.771350Z digest=sha256:16c35749b4462a1761f22d008c931d3bf0b2836ef3d5c1068006ad1b5f38084f

Observation 21a9c1d6-595c-4cb0-9df0-32949282d797 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.775855Z digest=sha256:307ba3aa9f885515343b19c40db9fb7346e394a8a061f7ca53f0d777ecef6c1d

Observation 2d5c8a7b-2e2e-4611-8957-ff5225150e0b · outbound

This paper cites The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.880239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.780004Z digest=sha256:63f49e8355897867a70d6208b0267326b9d016031c73d33133e9b003f8808113

Observation dbacfc98-69da-4b05-ac89-594e5b700859 · outbound

This paper cites Investigating the Limitations of Transformers with Simple Arithmetic Tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Investigating the Limitations of Transformers with Simple Arithmetic Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.784025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.784025Z digest=sha256:b374753b9550c8af8a2a6913d42777a23e9724f3f1019f37873a8f0ff71f6b63

Observation 0211457a-d6c4-426d-b890-b1f416dfa5b7 · outbound

This paper cites In-context Learning and Induction Heads.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer In-context Learning and Induction Heads

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.788567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.788567Z digest=sha256:f9303fa7209ae5a95ba1455ba6436018943fda5899da13c791a04c537efe895e

Observation 6a25774c-81d1-49bd-bc77-47473673ab53 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.793335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.793335Z digest=sha256:dfa5120b8658dc5987e8f277b3fea1f04fe25b6115c228c96fe804826dab0dfb

Observation 7dcd428a-9b5d-4e5f-987c-30bbdeab8320 · outbound

This paper cites Mechanistic Design and Scaling of Hybrid Architectures.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic Design and Scaling of Hybrid Architectures

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.798604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.798604Z digest=sha256:1dc1e26178b1275cb18c3a6be3e1a91427d3fc33b65d60d68e154893816793f1

Observation 021dc086-0b6c-4db4-8065-c4b2caa7fd2e · outbound

This paper cites Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.804007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.804007Z digest=sha256:c96eba210426991c80184c529ed9cc4e7a15150a822a4c221c7fafca2588a60d

Observation 5569d08c-f9ac-48ce-88af-7ae4d248fd97 · outbound

This paper cites Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.808829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.808829Z digest=sha256:282491877e23bb1604b8c1ec213de3e854609beeeea6faf0de2d037bbcd78e83

Observation dd623bde-0640-418a-83e6-52d2867d6da9 · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers, parallel computation, and logarithmic depth

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.814818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.814818Z digest=sha256:2d43cc33f37403eef1ab4a6d2935af2598535b3cd60ce9c47daf89456b7692ce

Observation 11d2c10c-c4bd-4ca0-97d0-bf2ca5d19531 · outbound

This paper cites Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.855012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.819921Z digest=sha256:ec6e466062eb65409a987c5c627e3ba9202936993aaff25f9756bf79ff44c82d

Observation 4aaf9a34-45f3-46e3-bf4c-199425e2134a · outbound

This paper cites Deep Information Propagation.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Information Propagation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.825593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.825593Z digest=sha256:27e8243c38822e5caa8d20608375dbd3fa9f055c5b353d824ee32eebf9923412

Observation 3eaae48b-ba2b-486d-a579-8517c8271da9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.830990Z digest=sha256:ca0a3e9523be738b4b7c8c695415787f6a442ceb2a872fd3cf85ecd4942c58fe

Observation a6875d19-b787-4d3f-a7f3-4611c9930bcd · outbound

This paper cites Synthesizer: Rethinking self-attention in transformer models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Synthesizer: Rethinking self-attention in transformer models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.827631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.836163Z digest=sha256:c27dcb81371d4001449c8f1e5f85d885cb8a4260e1c2b09e5b218706ccc53c59

Observation 63bf97f3-c7f4-429d-99be-880d433f5eb0 · outbound

This paper cites Efficient Transformers: A Survey.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient Transformers: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.842653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.842653Z digest=sha256:9f8bdcf8adefdb530715880526519081f47ce513db1865a7931422c8c3b60f6b

Observation b34f8d58-6c3e-4c0d-99b8-b42f47b9359d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.848934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.848934Z digest=sha256:31938394c5c0d60b0ee613b956189980d0fb1427bb75e9a2e92ee1891dcdadbd

Observation 0d52af51-62e5-487e-b502-cf2f62f01f67 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.853190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.853190Z digest=sha256:b84e6d6762b0e1a2f5636e5f2d2a2fd8a02089a1516b7a0301dd9eb5c3a70304

Observation dc020708-d608-41be-92e3-0f469d4e5e57 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.858670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.858670Z digest=sha256:1477faade804ff72c1614a460f0c39ca9f941f98e984616555584ee028333d4b

Observation d3defb31-0d36-4736-8d53-77758ef0d9be · outbound

This paper cites Efficient streaming language models with attention sinks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient streaming language models with attention sinks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.862990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.862990Z digest=sha256:0a6b2c54a84d42e986e0442cd66ed334a327e23c14df113586934cd83b6d50d3

Observation 096338b0-aec6-4fa4-8739-8d2059421669 · outbound

This paper cites On layer normalization in the transformer architecture.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer On layer normalization in the transformer architecture

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.867299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.867299Z digest=sha256:2640d83ed24da9ce6e35f80d3a044413c094320cef27ad1388bf8e95089d7e27

Observation e41da55c-0d29-4096-a942-585c53b48eb9 · outbound

This paper cites Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.782386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.872566Z digest=sha256:5339f7d02262d94487e8f0c8646bea91e98caa87736b53159faf3e711427e014

Observation 62a6dff9-fa4b-41a5-83b7-798521dfa5b2 · outbound

This paper cites Knowledge Circuits in Pretrained Transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge Circuits in Pretrained Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.877534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.877534Z digest=sha256:e0f33266a364294ce0d018afc3aade01b276e69082f5355700ba96a61158e7ea

Observation 04f9a6a8-d1a4-4ed3-a4b2-cb39ecf1b7a0 · outbound

This paper cites Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.768403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.882846Z digest=sha256:dbdfb02715e734207f59dc8257612545b932df5bc1e1ba66c3437fc4fd0eaca6

Observation 56cf4fe1-77de-43a9-a710-6a90ec70ae58 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.753867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.887654Z digest=sha256:42677515a0726190185d069f2750a5382cbfa6c4dfeb5d38a4476867a1d425c2

Observation 82e33590-6882-4b6a-9e7f-16ecdf82cd1a · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Understanding deep learning requires rethinking generalization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.891916Z digest=sha256:2aa01aaea394c5cd50d4a998ccf64c9ca92211160e3441666e5ccb371643beeb

Observation f6072648-2491-47d1-9f03-c2373e9c0d1c · outbound

This paper cites Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.011772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.896639Z digest=sha256:c49721b065e90ff31e052870ebfde897365039aa3d685948a8e4b783c6cff9f1

Observation 294b2fae-c881-45c7-a025-2371793ba2e9 · outbound

This paper cites Character-level Convolutional Networks for Text Classification.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Character-level Convolutional Networks for Text Classification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.901611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.901611Z digest=sha256:9ad92848889187d3d861aa3d4e3401d7077a71f2e278fba48e8ee1e1c10d5d33

Observation 8b1f9f34-206f-4e51-ae70-d8265d664326 · outbound

This paper cites Algorithmic capabilities of random transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Algorithmic capabilities of random transformers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.740198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:57:22.908318Z digest=sha256:a59ee593e9bf56c1163e71652b17eb3e19e6a213cec1d34bb773af168fcb8101

Pith citing papers

Observation 1a7b7cd7-c4d6-4a78-9f4e-10c11430db8b · inbound

Provable Knowledge Acquisition and Extraction in One-Layer Transformers cites this paper.

Provable Knowledge Acquisition and Extraction in One-Layer Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:10:44.852070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T23:05:56.687644Z digest=sha256:47619de4f1283a8fe85cfd6616bb1720efa4e469ffaf42ba6265a93b2b29b26f

Observation 25956d3d-2ece-4dbc-bcf9-d6d2a8a4b5f6 · inbound

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers cites this paper.

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:07.096284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:15:07.096284Z digest=sha256:b2aaf81eb4b5607d8df5af06fba3858c2b495f122f4ab256d7a8d1eab6c53ba7

Observation b9087432-cd4e-42a1-8b5b-3fa872b1f2c0 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:19.053632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:19:25.483640Z digest=sha256:87ee134710e35c7066121d824abec5ff706df5674ce05e86540f928167ae6f9c

Observation 5cd2c773-03ed-4a2b-a62c-a709c6b164d2 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:54:51.174733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T11:54:29.436149Z digest=sha256:d104cd48ee99dc77f008f2e91f7a8679b5b794530d5ede23c045a64e8bac8ab8

Observation 2ef29a84-3c61-4c99-a577-712726e622bb · inbound

Procedural Pretraining: Warming Up Language Models with Abstract Data cites this paper.

Procedural Pretraining: Warming Up Language Models with Abstract Data Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:55:10.047706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:55:10.047706Z digest=sha256:bd35ea71addf820d471ac12f26f1ec95947722458f821ec786c00ed034924cf6

Observation 5c3d438d-9afd-463d-93e2-61c7f1e8720e · inbound

Geometry-Calibrated Conformal Abstention for Language Models cites this paper.

Geometry-Calibrated Conformal Abstention for Language Models Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:28.657564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T05:56:03.031192Z digest=sha256:67d12f4662e595dd343b741c779a92014563c598d94ce052d854cb39fb69e2f7

Observation 3103343d-d627-4c1d-b9a7-94dba112716b · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.807728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:47:41.717389Z digest=sha256:d820c21e65a1d4f7f903aa8d7f8ae418b365170f7428741d370a482263e39a41

Observation b314efdc-8adb-44b0-8ae0-3031b2ecdab7 · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:15:11.538789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T07:15:02.755541Z digest=sha256:eabfbbcf540c499b607eebd665ac1b50036fe49ff34c18cfdbcce4e3903944b5

Observation 20037440-1bb0-4963-9551-f74d7b62b213 · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T11:17:06.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:17:06.440358Z digest=sha256:b03c3c655a810442a8e185c518d572381642c61eac5f76b85ef8f7cc07934c36

Observation 028dea82-b222-470e-bee6-0ef6955c05fe · inbound

Fixed Universal Transformers cites this paper.

Fixed Universal Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.217756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:24:18.984754Z digest=sha256:ff67b79a277eed8fa893e1d4bb764ad47f8b242f2c409d3c189802e9e7494fdc

Observation 8dbf2c27-2496-441f-bdef-32e954d60dcf · inbound

Activation-Based Active Learning for In-Context Learning: Challenges and Insights cites this paper.

Activation-Based Active Learning for In-Context Learning: Challenges and Insights Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.882218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T06:27:29.827899Z digest=sha256:620dcd4746fbf46b3244ad151a55efcf27ef12eb7aa057d4c5102b31e0235d53