Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:21.549139Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.06179.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:21.549139Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b186e2f3-97c6-49a0-869d-367bed3f96c2 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn to implement preconditioned gradient descent for in-context learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86348a3-4561-4195-b6ea-de5fbd1c5f02 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Block coordinate descent for neural networks provably finds global minima
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed9c1163-000a-4186-bcff-5fc84f90fc4f · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e07cc56c-df49-44ff-8f9e-95ae7a63a9e7 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Neural Machine Translation by Jointly Learning to Align and Translate
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10cf59d9-00e8-4344-8bf5-a1783691c633 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Ability and Limitations of Transformers to Recognize Formal Languages
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da02f76-adfb-40e5-a345-516b08524857 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Computational Power of Transformers and its Implications in Sequence Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e95b74c-6263-4093-bc2a-4b82c380afc1 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1e61cd4-1248-4fe4-9c20-fc4aaaf28195 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536d2bdd-a47a-4ba5-935b-198c6aa00d06 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Decision Transformer: Reinforcement Learning via Sequence Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56b1e1c8-31ae-4eb2-a06f-9c493c7e724e · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Provably learning a multi-head attention layer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a550f40-a8a8-4ee9-b9cd-7b0e0e2fd991 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Attention with Performers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd45c953-249f-42de-9f57-f256a71e4bd0 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Optimization and Generalization of Multi-head Attention
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9661c3a-1b2a-4d0c-a859-bb38637c1eae · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99701b9-2adf-497f-9573-589492dc5b38 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4f7717-7ac1-4bc6-b735-738760779656 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Inductive Biases and Variable Creation in Self-Attention Mechanisms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ce5238-6b0a-4348-94a3-7f4802a095c6 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A mathematical framework for transformer circuits
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ae237e-f46c-40aa-868c-c80d2b355632 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22bbf9eb-1c4c-4bc6-92ea-a01c9e52edb3 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization M., and Fan, J
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6362e547-8525-4924-bdcc-060998abbc7a · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On Limitation of Transformer for Learning HMMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8e3379-4fcd-4c62-a39f-9cf93381c4c3 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization In-context convergence of transformers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f2c682a-00e9-44c6-927e-990e382e8067 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a492e44c-f911-481a-84ba-313a963d527c · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Vision Transformers provably learn spatial structure
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bc12942-9cb9-4d32-8cc7-2ca58c640eff · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation daa10853-5dcd-416e-a272-2b2af1fbdbb7 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43d11c9-bdd9-46cc-ab9f-22c1271550ef · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9a66d0-e2b3-4bc2-8315-fd07ad268ec9 · outbound
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 690e44ed-757d-411e-868e-ae61c42ee29c · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9d517c-5ac8-4e1f-a088-3fbba6efc302 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization SimA: Simple Softmax-free Attention for Vision Transformers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e9af55a-da66-4b5e-b0a8-190a9c3b0978 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d5a435-75ec-4bac-ab9e-e3aad03ec028 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization The Closeness of In-Context Learning and Weight Shifting for Softmax Regression
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19f856e-914a-4962-bedf-4c7694a05efa · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Expressive Power of Self-Attention Matrices
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4de1ab7e-9de2-452b-9150-05f154133b75 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Learn Shortcuts to Automata
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a00c977-f4d6-4b39-9d34-8f98f198fe77 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Transformers in Solving POMDPs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 609770b0-6828-4366-8570-e6ee8d580761 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Your transformer may not be as powerful as you expect
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eccb2b48-c5b2-4ff5-abaa-3473c20fd9ba · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Expressive, But Are They Expressive Enough for Regression?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e29a8a8-5a31-4110-b43b-ed1dbccdd87e · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Representational Strengths and Limitations of Transformers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ee5037-49bf-4eaa-844f-c146f8fa6315 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db8e1eb0-9a49-4aa4-8bee-0609d5641170 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unraveling the gradient descent dynamics of transformers
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67c703c1-8a6b-4f82-be50-7ba0e106a85f · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b50d96b-e00c-4ba8-9ca7-4a4b3b64e704 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers as Support Vector Machines
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4bcdc5-53d4-47c7-a369-99fd34080d22 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2bc851c-fdb9-4894-8e9b-b6dbb3bdb473 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Introduction to Matrix Concentration Inequalities
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d115a87-ac16-4df7-9565-0ad748ae8882 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Attention Is All You Need
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0005a2-a080-41ca-ba0c-86ce5636c97f · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linformer: Self-Attention with Linear Complexity
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30eb5080-6990-49a8-882f-2efd1e737368 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4559da0e-1c67-4936-926e-00f186f34955 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c471eee-41ae-4717-a298-d15b0b9f958f · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65e87cbc-8b73-4884-8d06-560981f90fff · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Self-Attention Networks Can Process Bounded Hierarchical Languages
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e11c36a9-ba77-42ee-894a-2d3458285132 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 811b1712-1d5b-43dd-8bef-2b8dfe8355f3 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Global Convergence of Block Coordinate Descent in Deep Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1374c329-c8d2-4784-9261-e9b11af7be25 · outbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Efficient Compilers, Provably
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.