Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:22.908318Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 11 inbound Pith citation observations for arXiv:2506.01115.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:22.908318Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.096284Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T07:56:47.880747Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8363ba87-329c-42c9-9ecc-0a23659bea68 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9abca2d-5c2b-440f-9692-f09073b1fa79 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The Curious Case of Benign Memorization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73175888-b0e0-4f1b-a063-df84c3ac5ce3 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang.On exact computation with an infinitely wide neural net
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e41ea57a-c370-4585-bf0c-5511b8b9b17a · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A closer look at memorization in deep networks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 117151e8-5cca-4bda-b64b-6e54be0079b1 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Scaling mlps: A tale of inductive bias
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 879bb15d-bee3-410b-a93f-0189690761f1 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic interpretability for AI safety - a review.Transactions on Machine Learning Research, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea2cbf4-4bb5-4b85-af8e-f660e06859a4 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Birth of a transformer: A memory viewpoint.Advances in Neural Information Processing Systems, 36:1560–1588, 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 662cb06e-280c-496c-b623-4dc346d22452 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3191c054-6843-42be-a1a1-84be8dbd1f98 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers generalize differently from information stored in context vs in weights
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c7a563-58a0-4209-ae7e-594a24231996 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Distributional associations vs in-context reasoning: A study of feed-forward and attention layers.ICLR, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68fa1850-9619-4b11-837e-0cf11690ca00 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge localization: Mission not accomplished? enter query localization! InProceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ca9d54d-76bd-4019-af87-fba6bb10bc3e · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Summing up the facts: Additive mechanisms behind factual recall in llms, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e846ff2-6ed8-4deb-8034-b1fb20869d90 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Induction heads as an essential mechanism for pattern matching in in-context learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83408a97-f0ab-447b-a5e6-c01f0c6a15c1 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge neurons in pretrained transformers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cc6a85-ddda-418d-9e36-2645fb4adb39 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28cc1dae-114a-43c7-a596-2b6d94072207 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d511882-7077-4c50-b2ab-5aa70f144bfa · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Edelman, eran malach, and Surbhi Goel
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6de929e4-298d-4f9f-bd46-dcf17694bbd6 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58bf9a70-406d-4f25-8238-7eedfc9358da · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer Feed-Forward Layers Are Key-Value Memories
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47962c0-f560-41db-b6ee-fa5f74c6d609 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers are key-value memories
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f87075-542c-4a28-afc8-718b2ecba1d3 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc95dbb-4299-4414-a852-483b5d696015 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Dissectingrecalloffactualassociations inauto-regressivelanguagemodels
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e0fa06-aee0-48fa-b920-79811df9a4f2 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9631f467-18cd-46a7-aa33-fb364dfc04cf · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Smith, and Roy Schwartz
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2849f7ac-531f-4d2e-9980-80ffde91facd · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Simplifying Transformer Blocks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20848180-274a-4d69-a981-68c6b2675f35 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 547444a1-677d-4f56-a61f-815b88a3eef3 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer What is the best multi-stage architecture for object recognition? In2009 IEEE 12th International Conference on Computer Vision, pages 2146–2153, 2009
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657d772e-82b4-4721-82b6-807cdc2c3692 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Lexico: Extreme kv cache compression via sparse coding over universal dictionaries, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1487f04a-6d72-4d38-87fd-03cc49bd03a7 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Neural Networks as Gaussian Processes
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712c06fc-030a-4711-a0f1-dbb80d30ae53 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer FNet: Mixing Tokens with Fourier Transforms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2b73b7-af45-463d-9a16-cae3f6618d03 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795–10808, 2022
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59b06326-5916-47b7-9a17-d49c8b382d79 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8d18586-9ccc-4f75-92cb-096e3db23870 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating and Editing Factual Associations in GPT
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0aca01b-a1f7-4fd7-a5da-eabfdb3d31f5 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Pointer Sentinel Mixture Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1565742c-8d44-47e6-a5d7-46ce1da0bc03 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Language models implement simple word2vec-style vector arithmetic, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80ff018a-334f-4a56-abb3-99e13bcdbd6a · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Universal approximation property of random neural networks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21a9c1d6-595c-4cb0-9df0-32949282d797 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.Advances in Neural Information Processing Systems, 35:27198–27211, 2022
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5c8a7b-2e2e-4611-8957-ff5225150e0b · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The shaped transformer: Attention models in the infinite depth-and-width limit.Advances in Neural Information Processing Systems, 36:54250–54281, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbacfc98-69da-4b05-ac89-594e5b700859 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Investigating the Limitations of Transformers with Simple Arithmetic Tasks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0211457a-d6c4-426d-b890-b1f416dfa5b7 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer In-context Learning and Induction Heads
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a25774c-81d1-49bd-bc77-47473673ab53 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dcd428a-9b5d-4e5f-987c-30bbdeab8320 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mechanistic Design and Scaling of Hybrid Architectures
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021dc086-0b6c-4db4-8065-c4b2caa7fd2e · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Exponential expressivity in deep neural networks through transient chaos.Advances in neural information processing systems, 29, 2016
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5569d08c-f9ac-48ce-88af-7ae4d248fd97 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd623bde-0640-418a-83e6-52d2867d6da9 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Transformers, parallel computation, and logarithmic depth
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11d2c10c-c4bd-4ca0-97d0-bf2ca5d19531 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4aaf9a34-45f3-46e3-bf4c-199425e2134a · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Information Propagation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eaae48b-ba2b-486d-a579-8517c8271da9 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6875d19-b787-4d3f-a7f3-4611c9930bcd · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Synthesizer: Rethinking self-attention in transformer models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63bf97f3-c7f4-429d-99be-880d433f5eb0 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient Transformers: A Survey
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34f8d58-6c3e-4c0d-99b8-b42f47b9359d · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Gemini: A Family of Highly Capable Multimodal Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d52af51-62e5-487e-b502-cf2f62f01f67 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer LLaMA: Open and Efficient Foundation Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc020708-d608-41be-92e3-0f469d4e5e57 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3defb31-0d36-4736-8d53-77758ef0d9be · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient streaming language models with attention sinks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096338b0-aec6-4fa4-8739-8d2059421669 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer On layer normalization in the transformer architecture
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41da55c-0d29-4096-a942-585c53b48eb9 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Mean field residual networks: On the edge of chaos.Advances in neural information processing systems, 30, 2017
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62a6dff9-fa4b-41a5-83b7-798521dfa5b2 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Knowledge Circuits in Pretrained Transformers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f9a6a8-d1a4-4ed3-a4b2-cb39ecf1b7a0 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Locating factual knowledge in large language models: Exploring the residual stream and analyzing subvalues in vocabulary space, 2024
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56cf4fe1-77de-43a9-a710-6a90ec70ae58 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Are transformers universal approximators of sequence-to-sequence functions? InInternational Conference on Learning Representations, 2020
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82e33590-6882-4b6a-9e7f-16ecdf82cd1a · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Understanding deep learning requires rethinking generalization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6072648-2491-47d1-9f03-c2373e9c0d1c · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 294b2fae-c881-45c7-a025-2371793ba2e9 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Character-level Convolutional Networks for Text Classification
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1f9f34-206f-4e51-ae70-d8265d664326 · outbound
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Algorithmic capabilities of random transformers
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a7b7cd7-c4d6-4a78-9f4e-10c11430db8b · inbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25956d3d-2ece-4dbc-bcf9-d6d2a8a4b5f6 · inbound
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9087432-cd4e-42a1-8b5b-3fa872b1f2c0 · inbound
Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cd2c773-03ed-4a2b-a62c-a709c6b164d2 · inbound
Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ef29a84-3c61-4c99-a577-712726e622bb · inbound
Procedural Pretraining: Warming Up Language Models with Abstract Data Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3d438d-9afd-463d-93e2-61c7f1e8720e · inbound
Geometry-Calibrated Conformal Abstention for Language Models Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3103343d-d627-4c1d-b9a7-94dba112716b · inbound
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b314efdc-8adb-44b0-8ae0-3031b2ecdab7 · inbound
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20037440-1bb0-4963-9551-f74d7b62b213 · inbound
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028dea82-b222-470e-bee6-0ef6955c05fe · inbound
Fixed Universal Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8dbf2c27-2496-441f-bdef-32e954d60dcf · inbound
Activation-Based Active Learning for In-Context Learning: Challenges and Insights Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.