Pith. sign in

Paper Citation Record · LEDGER

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

As of 10 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 11 inbound Pith citation observations for arXiv:2506.01115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01115 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:22.908318Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.096284Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.880747Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8363ba87-329c-42c9-9ecc-0a23659bea68 · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.571092Z digest=sha256:8eadeb10d2e8df0a744daca4ab73d709bfed12192f4bc9079022b136b5eed7a8

Observation e9abca2d-5c2b-440f-9692-f09073b1fa79 · outbound

This paper cites The Curious Case of Benign Memorization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The Curious Case of Benign Memorization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.702593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.577524Z digest=sha256:e2e142334ea8677c242d6b29360d6fd9c0305d837728ccf34379bdc911ab6ab6

Observation 73175888-b0e0-4f1b-a063-df84c3ac5ce3 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.158136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.583250Z digest=sha256:87d26bbc04e516ee3ca4b8e308e90554b2927164a177e4d567649f24eee03427

Observation e41ea57a-c370-4585-bf0c-5511b8b9b17a · outbound

This paper cites A closer look at memorization in deep networks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A closer look at memorization in deep networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.143410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.588325Z digest=sha256:fd894f721368ea04dc3b438eb6e6861ed10e307244f51008d1fcfe5124f10317

Observation 117151e8-5cca-4bda-b64b-6e54be0079b1 · outbound

This paper cites Scaling mlps: A tale of inductive bias.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Scaling mlps: A tale of inductive bias

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.123851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.593277Z digest=sha256:f11015300b4db8946bf9b92368d49ebd873f53316a562e4baee9992538f688fd

Observation 879bb15d-bee3-410b-a93f-0189690761f1 · outbound

This paper cites Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.599011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.599011Z digest=sha256:d71a28372a5e70a410f09a59ee7629272ad4c5034f6247937d1a158b167fd3c2

Observation 0ea2cbf4-4bb5-4b85-af8e-f660e06859a4 · outbound

This paper cites Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.096487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.605192Z digest=sha256:341f3378108c18ab81478d6b1117a4fa1522c31f893e455286dd2fb548af916a

Observation 662cb06e-280c-496c-b623-4dc346d22452 · outbound

This paper cites Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.631525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.611391Z digest=sha256:a2bb9544a55f824fa529882213067078e018cffe4d66059a016991eade74d936

Observation 3191c054-6843-42be-a1a1-84be8dbd1f98 · outbound

This paper cites Transformers generalize differently from information stored in context vs in weights.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers generalize differently from information stored in context vs in weights

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.619625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.619625Z digest=sha256:010b3e40f9db6b01c3cd66191c533694cb0bd3154182432e678b010135ec053c

Observation 30c7a563-58a0-4209-ae7e-594a24231996 · outbound

This paper cites Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.082310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.626924Z digest=sha256:3c44ad949eec0107ca6751e9ba613123bada5cb5746711bda91f76502a218db6

Observation 68fa1850-9619-4b11-837e-0cf11690ca00 · outbound

This paper cites Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.068319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.633098Z digest=sha256:a39e8f2d7ef449ee997d332a2d5ea72de29c2d62063910a668e0461639402b7d

Observation 8ca9d54d-76bd-4019-af87-fba6bb10bc3e · outbound

This paper cites Summing up the facts: Additive mechanisms behind factual recall in llms, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Summing up the facts: Additive mechanisms behind factual recall in llms, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.054015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.639502Z digest=sha256:a2f59163c663fbb27819d793a6ce4026dccce4b8d52e81cef51e3d68986ad4fe

Observation 1e846ff2-6ed8-4deb-8034-b1fb20869d90 · outbound

This paper cites Induction heads as an essential mechanism for pattern matching in in-context learning.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Induction heads as an essential mechanism for pattern matching in in-context learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.040339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.647499Z digest=sha256:c3312790698ea46f2600cb28db91222a5af2a88907b40f147dd4019e915badcc

Observation 83408a97-f0ab-447b-a5e6-c01f0c6a15c1 · outbound

This paper cites Knowledge neurons in pretrained transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge neurons in pretrained transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.655162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.655162Z digest=sha256:3286cce7bbe10bc2a2dcf4c37180b41f6555448bed078574861a24dfd31bf954

Observation 44cc6a85-ddda-418d-9e36-2645fb4adb39 · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.027077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.661558Z digest=sha256:472e7d4ad2c0feca720bbbfc7252445ad206bab475c94eaf3ea0690f570925f0

Observation 28cc1dae-114a-43c7-a596-2b6d94072207 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.667162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.667162Z digest=sha256:53cfcb224919134cb2c87427501550f78da64be9b3101f65246acb22a728bfc0

Observation 1d511882-7077-4c50-b2ab-5aa70f144bfa · outbound

This paper cites Edelman, eran malach, and Surbhi Goel.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Edelman, eran malach, and Surbhi Goel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:24.008710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.672227Z digest=sha256:6d4eccbb4f5d6e4e16b7378459c1e7603ac602be518a40c1e81188889924ad7f

Observation 6de929e4-298d-4f9f-bd46-dcf17694bbd6 · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.989975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.677623Z digest=sha256:2713b480b8950084dd63b04f32666c30b69386cd5f3633726e80c13e66556623

Observation 58bf9a70-406d-4f25-8238-7eedfc9358da · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer Feed-Forward Layers Are Key-Value Memories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.682505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.682505Z digest=sha256:29e8f403b6edf72852c13b08f5f2f7de723f196dc5dd4caa53ba121b593c0e96

Observation c47962c0-f560-41db-b6ee-fa5f74c6d609 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers are key-value memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.687732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.687732Z digest=sha256:6a1f5e620aa8b4c345969e7832f9b6119aecf94d71a2dc0c2e519ecb86ae0423

Observation c0f87075-542c-4a28-afc8-718b2ecba1d3 · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.692163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.692163Z digest=sha256:cf265f636f744c49d1c4893d4ec97126626fa5800248bb6b7debf063b973440f

Observation 2dc95dbb-4299-4414-a852-483b5d696015 · outbound

This paper cites Dissectingrecalloffactualassociations inauto-regressivelanguagemodels.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Dissectingrecalloffactualassociations inauto-regressivelanguagemodels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.696572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.696572Z digest=sha256:386ef83983483200ce489d35bbcfe3ab2bdbb94bba30cb8c2d3fdc379d51a806

Observation a3e0fa06-aee0-48fa-b920-79811df9a4f2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.701367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.701367Z digest=sha256:cb7b3d6c60c032cdb98bd1a8ce744e43f42a127ccc5f00fa8df6381361530049

Observation 9631f467-18cd-46a7-aa33-fb364dfc04cf · outbound

This paper cites Smith, and Roy Schwartz.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Smith, and Roy Schwartz

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.974388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.705870Z digest=sha256:12c213673a55e2d5f14b6b0744508b761e46ffd28f0be1f76e0d4ac510ceffb7

Observation 2849f7ac-531f-4d2e-9980-80ffde91facd · outbound

This paper cites Simplifying Transformer Blocks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Simplifying Transformer Blocks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.710780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.710780Z digest=sha256:fd63527522461242868c5d4b960bd6fd189a74e04b6b3721468d43de95d4ed40

Observation 20848180-274a-4d69-a981-68c6b2675f35 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.716704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.716704Z digest=sha256:5312ac300a9e68870ad66b0304549e9c0a11d4cbdc3e76370bfaad8d25330cb2

Observation 547444a1-677d-4f56-a61f-815b88a3eef3 · outbound

This paper cites What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.722380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.722380Z digest=sha256:b12adfeff3da4f9252afaf195cec0656857eed4180161da5219da0c2085a4f3b

Observation 657d772e-82b4-4721-82b6-807cdc2c3692 · outbound

This paper cites Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.960220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.728573Z digest=sha256:34ab27bf17de08768fe4afaedcfaa2bc37ceec7a781a5ea42c7acc98341ffb26

Observation 1487f04a-6d72-4d38-87fd-03cc49bd03a7 · outbound

This paper cites Deep Neural Networks as Gaussian Processes.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Neural Networks as Gaussian Processes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.735092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.735092Z digest=sha256:01a7794fa23c91a672753b8b0fbae6fd5d879d27a9555aea6220c8c398ea0ed4

Observation 712c06fc-030a-4711-a0f1-dbb80d30ae53 · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer FNet: Mixing Tokens with Fourier Transforms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.740900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.740900Z digest=sha256:b2faae42e17dd1afee0413ba8e136888d7be3c4830134d3ead86ecdd257c285a

Observation 5c2b73b7-af45-463d-9a16-cae3f6618d03 · outbound

This paper cites The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.946050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.745978Z digest=sha256:4bd13cac85f08fd0c59f6d10c2d7d607a9b858853c1f2212c4c105072898aa72

Observation 59b06326-5916-47b7-9a17-d49c8b382d79 · outbound

This paper cites Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.750524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.750524Z digest=sha256:d4ff6276d82e0e61a67828504570af5a9a3efcd977632ffc0e7c2c54512f90ce

Observation f8d18586-9ccc-4f75-92cb-096e3db23870 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating and Editing Factual Associations in GPT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.755732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.755732Z digest=sha256:ba39c2e4fdc301f404872bfea4358906b884ce1808ca7de7fa78f719323eeff7

Observation a0aca01b-a1f7-4fd7-a5da-eabfdb3d31f5 · outbound

This paper cites Pointer Sentinel Mixture Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Pointer Sentinel Mixture Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.760917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.760917Z digest=sha256:f37c8ebde10124364b9192097fa6217065314a768833c39e1b2c928116cf1dc2

Observation 1565742c-8d44-47e6-a5d7-46ce1da0bc03 · outbound

This paper cites Language models implement simple word2vec-style vector arithmetic, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Language models implement simple word2vec-style vector arithmetic, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.927859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.765913Z digest=sha256:502c7f0b6331ef1017a604ad7601d4c7e31bdd3d10f27689fc6ffbc2aa8b1158

Observation 80ff018a-334f-4a56-abb3-99e13bcdbd6a · outbound

This paper cites Universal approximation property of random neural networks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Universal approximation property of random neural networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.911077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.771350Z digest=sha256:19de6a817c5b2277e8766935e56f942e61523e2417847295a7e96e9486bf1b53

Observation 21a9c1d6-595c-4cb0-9df0-32949282d797 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.775855Z digest=sha256:6fa1c357ff53c92c31e38beaec8a50476a84abfdd1336604c01b811dc5b8a561

Observation 2d5c8a7b-2e2e-4611-8957-ff5225150e0b · outbound

This paper cites The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.880239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.780004Z digest=sha256:401f85e3aec689e531623ba6a449a5f08fb91fd777c40d8fb69d8f7fea67d712

Observation dbacfc98-69da-4b05-ac89-594e5b700859 · outbound

This paper cites Investigating the Limitations of Transformers with Simple Arithmetic Tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Investigating the Limitations of Transformers with Simple Arithmetic Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.784025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.784025Z digest=sha256:e01344d1a49b7eeee01783da7be95d2a49c47ce2b28280310c8b19ef6c9e1fa9

Observation 0211457a-d6c4-426d-b890-b1f416dfa5b7 · outbound

This paper cites In-context Learning and Induction Heads.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer In-context Learning and Induction Heads

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.788567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.788567Z digest=sha256:f99a8229156f7b651a3bd8cf633f464eb56a23c20a92eade405067b7cd8df0b3

Observation 6a25774c-81d1-49bd-bc77-47473673ab53 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.793335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.793335Z digest=sha256:48b92c6bb8929bde2376f21e23367ef824b49d066fe8d43717f7df8444419298

Observation 7dcd428a-9b5d-4e5f-987c-30bbdeab8320 · outbound

This paper cites Mechanistic Design and Scaling of Hybrid Architectures.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic Design and Scaling of Hybrid Architectures

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.798604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.798604Z digest=sha256:05c01a7485c9a40bae78e794a5a3e50608f2faee228c32c919e73f1091a32812

Observation 021dc086-0b6c-4db4-8065-c4b2caa7fd2e · outbound

This paper cites Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.804007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.804007Z digest=sha256:ee57be00baf7f340712208ba88083a4c297e7bf79f2da04daf3b9a6eaf16a09a

Observation 5569d08c-f9ac-48ce-88af-7ae4d248fd97 · outbound

This paper cites Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.808829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.808829Z digest=sha256:99de8bf833198240c4b8f789da9397d5e292e0d39b479ff5d341a4c1ab0cf6b6

Observation dd623bde-0640-418a-83e6-52d2867d6da9 · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers, parallel computation, and logarithmic depth

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.814818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.814818Z digest=sha256:efd7fb62000f840ea2a193691636037e63d15c0ddf05e2d479b0d9b849145a4f

Observation 11d2c10c-c4bd-4ca0-97d0-bf2ca5d19531 · outbound

This paper cites Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.855012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.819921Z digest=sha256:8cdb585d5c599874cb496e197261277b35ddea45029d84859ae7ff622a939fd5

Observation 4aaf9a34-45f3-46e3-bf4c-199425e2134a · outbound

This paper cites Deep Information Propagation.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Information Propagation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.825593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.825593Z digest=sha256:0f9864854bb27d0d6922e254706eed2533d51f77aa37ebb671c2ac3b87af050e

Observation 3eaae48b-ba2b-486d-a579-8517c8271da9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.830990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.830990Z digest=sha256:0d85a8e79298c931761179bc248ff01935ee1e9d5e4d811c20857ec231f25c1c

Observation a6875d19-b787-4d3f-a7f3-4611c9930bcd · outbound

This paper cites Synthesizer: Rethinking self-attention in transformer models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Synthesizer: Rethinking self-attention in transformer models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.827631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.836163Z digest=sha256:2913d79d88cee811f902daf9bc4536685801d8dd59fd4c47277a194200edb148

Observation 63bf97f3-c7f4-429d-99be-880d433f5eb0 · outbound

This paper cites Efficient Transformers: A Survey.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient Transformers: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.842653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.842653Z digest=sha256:608279f83c93ca626dd130fda8c8d289a2f478e874f5b49b75352026e73e4be2

Observation b34f8d58-6c3e-4c0d-99b8-b42f47b9359d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.848934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.848934Z digest=sha256:8d3e4f538466aad7dd9d16747bc69136fa0835e567522bdc25b3bef5b4245b8d

Observation 0d52af51-62e5-487e-b502-cf2f62f01f67 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.853190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.853190Z digest=sha256:552bd06eccc227220538be2c00e87776d3ec5fee2aa754ec8e032c171dda529a

Observation dc020708-d608-41be-92e3-0f469d4e5e57 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.858670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.858670Z digest=sha256:939e837c32d772827b5e655bad65439e69edd4f8f5a9f6e433ddcc4d91012ae6

Observation d3defb31-0d36-4736-8d53-77758ef0d9be · outbound

This paper cites Efficient streaming language models with attention sinks.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient streaming language models with attention sinks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.862990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.862990Z digest=sha256:0fb17909cb90b7a27e3176059d2d86663ee0e83150874cb5e23174a29f1454e4

Observation 096338b0-aec6-4fa4-8739-8d2059421669 · outbound

This paper cites On layer normalization in the transformer architecture.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer On layer normalization in the transformer architecture

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.867299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.867299Z digest=sha256:fb566ac10ddf3071f355e95e52b38c85a4a742e878b40480f458c787cbdb7d7e

Observation e41da55c-0d29-4096-a942-585c53b48eb9 · outbound

This paper cites Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.782386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.872566Z digest=sha256:5b7024ce38b5cac25ae217e2cf70bd8019a248d1963c66c9835800da391e9596

Observation 62a6dff9-fa4b-41a5-83b7-798521dfa5b2 · outbound

This paper cites Knowledge Circuits in Pretrained Transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge Circuits in Pretrained Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.877534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.877534Z digest=sha256:6fbc882635dcd11ebf038b5bc4244808df3ff6e6ee2937a38383205e7688399b

Observation 04f9a6a8-d1a4-4ed3-a4b2-cb39ecf1b7a0 · outbound

This paper cites Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.768403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.882846Z digest=sha256:aaa65c693fd3c600d0f8572c2b9182c7d9a7eb9702865c8c8aaf10fd519c7c48

Observation 56cf4fe1-77de-43a9-a710-6a90ec70ae58 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.753867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.887654Z digest=sha256:e8f9a957aa93a224d71515a6285ccaac07ff906bf8ef518d7c2d1f6af2c10f56

Observation 82e33590-6882-4b6a-9e7f-16ecdf82cd1a · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Understanding deep learning requires rethinking generalization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.891916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.891916Z digest=sha256:61030bf41e5e34a2cd9a1250af09f8f25ec5bc528dc7840fab96ef925a8647d4

Observation f6072648-2491-47d1-9f03-c2373e9c0d1c · outbound

This paper cites Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:23.011772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.896639Z digest=sha256:708474a9d95411a5620c3c2c019497c61e02ffb27ce5900e7dfea3f93c0c1192

Observation 294b2fae-c881-45c7-a025-2371793ba2e9 · outbound

This paper cites Character-level Convolutional Networks for Text Classification.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Character-level Convolutional Networks for Text Classification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.901611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.901611Z digest=sha256:47545ede1243e63b8eac91967a6897cc45a294a5857e6c1d65fdeb0161b8b39c

Observation 8b1f9f34-206f-4e51-ae70-d8265d664326 · outbound

This paper cites Algorithmic capabilities of random transformers.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Algorithmic capabilities of random transformers

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:23.740198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:57:22.908318Z digest=sha256:155dd2f7987fb9e0806f0edd926980e171362b14e104a79691f78d1921b2c525

Pith citing papers

Observation 1a7b7cd7-c4d6-4a78-9f4e-10c11430db8b · inbound

Provable Knowledge Acquisition and Extraction in One-Layer Transformers cites this paper.

Provable Knowledge Acquisition and Extraction in One-Layer Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:10:44.852070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T23:05:56.687644Z digest=sha256:41be08ba4bfd37c22d0b5ec96d2b0b8fd6cb5d2cf86fdef71141e5fc042f5a97

Observation 25956d3d-2ece-4dbc-bcf9-d6d2a8a4b5f6 · inbound

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers cites this paper.

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:07.096284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:15:07.096284Z digest=sha256:57482ea64683eee3033f4550dae55767d5d66040da73644bb8a9cebd3270361e

Observation b9087432-cd4e-42a1-8b5b-3fa872b1f2c0 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:19.053632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:19:25.483640Z digest=sha256:1cd4e6c060257cdfe3e9d27bb03265aa0839ac930ca7a0dda80286ae74990647

Observation 5cd2c773-03ed-4a2b-a62c-a709c6b164d2 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:54:51.174733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:54:29.436149Z digest=sha256:58b2acd74e14bb71bc0d848577f8e6d362a3ccbc6e41490335b6a21d67a1c286

Observation 2ef29a84-3c61-4c99-a577-712726e622bb · inbound

Procedural Pretraining: Warming Up Language Models with Abstract Data cites this paper.

Procedural Pretraining: Warming Up Language Models with Abstract Data Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:55:10.047706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:55:10.047706Z digest=sha256:53d177aaf594cae4ce2d118f127c4a87f461cb72c4b2be57aeb056977910537c

Observation 5c3d438d-9afd-463d-93e2-61c7f1e8720e · inbound

Geometry-Calibrated Conformal Abstention for Language Models cites this paper.

Geometry-Calibrated Conformal Abstention for Language Models Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:28.657564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T05:56:03.031192Z digest=sha256:949184f678fcd1f4d1ef564f50ecc8de37d852bd540910f4eff65b61a6c33d74

Observation 3103343d-d627-4c1d-b9a7-94dba112716b · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.807728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T11:47:41.717389Z digest=sha256:893717597dbf77c93420341ef620b14b7429006ea95c9f3f339fa66841a7b9e0

Observation b314efdc-8adb-44b0-8ae0-3031b2ecdab7 · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:15:11.538789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:15:02.755541Z digest=sha256:bf5a24440cde2129d5bc1d6b303986b40f168a90ae59dedf926b8b45ae9f66a4

Observation 20037440-1bb0-4963-9551-f74d7b62b213 · inbound

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination cites this paper.

Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T11:17:06.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:17:06.440358Z digest=sha256:35e87eb8b1a8e9df02aa7c3a5232348c93459028a2c2d253706e7eef2a9454a0

Observation 028dea82-b222-470e-bee6-0ef6955c05fe · inbound

Fixed Universal Transformers cites this paper.

Fixed Universal Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.217756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:24:18.984754Z digest=sha256:2960b5922bd701d3fe322b24f8e7dc552c8b0e37de126a65fdcb74f037789210

Observation 8dbf2c27-2496-441f-bdef-32e954d60dcf · inbound

Activation-Based Active Learning for In-Context Learning: Challenges and Insights cites this paper.

Activation-Based Active Learning for In-Context Learning: Challenges and Insights Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.882218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T06:27:29.827899Z digest=sha256:b8680eaa29c4b0a4240084d409f0db58ffdf1a9777a12a25a6998ab7b3508cf5