Pith. sign in

Paper Citation Record · LEDGER

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization

As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.06179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06179 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:21.549139Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact8
  • verified fuzzy7
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b186e2f3-97c6-49a0-869d-367bed3f96c2 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.045295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.045295Z digest=sha256:06791333201ba4512c881c43045bd0bacbd9de2bf5b6a6022413d63336acdfb0

Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.118380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.118380Z digest=sha256:fa45afb49a71439a0f25b51be957215456269ef2949991a2243d6ada01a37351

Observation b86348a3-4561-4195-b6ea-de5fbd1c5f02 · outbound

This paper cites Block coordinate descent for neural networks provably finds global minima.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Block coordinate descent for neural networks provably finds global minima

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.632599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.308487Z digest=sha256:9c6d63a0889ae61473ed102c63135d6d3e0abf6e357fa2f96efb0c781618aa12

Observation ed9c1163-000a-4186-bcff-5fc84f90fc4f · outbound

This paper cites How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.430307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.430307Z digest=sha256:c029b1283bcc61fde2eca8587ad90d8cb6f9328978c31aba70140aa7ce291eaf

Observation e07cc56c-df49-44ff-8f9e-95ae7a63a9e7 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.520290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.520290Z digest=sha256:0f7d1f609e5194c4679f8b6dd40e982a073d999505918d9856f294b8dd7e343f

Observation 10cf59d9-00e8-4344-8bf5-a1783691c633 · outbound

This paper cites On the Ability and Limitations of Transformers to Recognize Formal Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Ability and Limitations of Transformers to Recognize Formal Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.604932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.604932Z digest=sha256:7b736b904b970dc7fb3b9a229b390e0758b42eb1b09798e33d501d4df51b60fd

Observation 3da02f76-adfb-40e5-a345-516b08524857 · outbound

This paper cites On the Computational Power of Transformers and its Implications in Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Computational Power of Transformers and its Implications in Sequence Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.691808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.691808Z digest=sha256:697f6d149ee03eaa43776fd6a71337c5028a03d9444287887e51b40cf017c0d9

Observation 7e95b74c-6263-4093-bc2a-4b82c380afc1 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:10:25.347936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.756489Z digest=sha256:818f78513065969e637f45a5c3e4d283e1863b5c23e8094287d834de52100beb

Observation d1e61cd4-1248-4fe4-9c20-fc4aaaf28195 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.825748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.825748Z digest=sha256:63c6fbf24b6f3482c8ce97d22837c65b99fae971966c36eea9baa91b4154b14c

Observation 536d2bdd-a47a-4ba5-935b-198c6aa00d06 · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.886542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.886542Z digest=sha256:4f086a699a6d76c02ee5c1738f8c4b620b10b0fd9942e9433a3d5e110cd98370

Observation 56b1e1c8-31ae-4eb2-a06f-9c493c7e724e · outbound

This paper cites Provably learning a multi-head attention layer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Provably learning a multi-head attention layer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.973138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.973138Z digest=sha256:a1a32319eb18d0259222973d92bd969e7d54e9809acf9ad4fa7818887a82824c

Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.068890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.068890Z digest=sha256:2a0baa3965898a0c96f34b9170dcbab70636004721782f74411fb24851127233

Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · outbound

This paper cites Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.138258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.138258Z digest=sha256:cff12ada3935ec0d83628ad8c0002f6241d02daac28e83263756710c6811a7ec

Observation 6a550f40-a8a8-4ee9-b9cd-7b0e0e2fd991 · outbound

This paper cites Rethinking Attention with Performers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Attention with Performers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.231632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.231632Z digest=sha256:7e1211a44c69d68628e75203decf7b04e1b1592050544c4526d84ae37167635b

Observation cd45c953-249f-42de-9f57-f256a71e4bd0 · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Optimization and Generalization of Multi-head Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.345237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.345237Z digest=sha256:2eb6eee3332b69ef00ca9ea9d92cae1d1aa3fe67984843ee6e36632b70971289

Observation d9661c3a-1b2a-4d0c-a859-bb38637c1eae · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.408915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.408915Z digest=sha256:6b9070d3931aa8c5b56b9dfd61cb0fdaddabd280edc4dc36b2ce6497885b4394

Observation f99701b9-2adf-497f-9573-589492dc5b38 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.511563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.511563Z digest=sha256:9bb8eca4e6ebdd18e779b7364ed4958354ceafabcc84b1a1935076554b086cb7

Observation 4d4f7717-7ac1-4bc6-b735-738760779656 · outbound

This paper cites Inductive Biases and Variable Creation in Self-Attention Mechanisms.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Inductive Biases and Variable Creation in Self-Attention Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.582759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.582759Z digest=sha256:f080575bb0b214816816a10121bd436d35074d11ae40234717c93ed6a4bff78c

Observation 88ce5238-6b0a-4348-94a3-7f4802a095c6 · outbound

This paper cites A mathematical framework for transformer circuits.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A mathematical framework for transformer circuits

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.672307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.672307Z digest=sha256:b04662091420b86a98f7c649c903d4cf3efb78d086336c48318cc2628dd39f53

Observation 31ae237e-f46c-40aa-868c-c80d2b355632 · outbound

This paper cites Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T06:10:22.175983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.767563Z digest=sha256:75d6eb0ab69f76d3819f76d8dd45a0a28dbadda28010af44e5b851101e1d9f6a

Observation 22bbf9eb-1c4c-4bc6-92ea-a01c9e52edb3 · outbound

This paper cites M., and Fan, J.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization M., and Fan, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.117893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.829721Z digest=sha256:f5355f5cc1ba5a789bfcc3a445ac98e6badde07d10502649029e5d2e4f007afa

Observation 6362e547-8525-4924-bdcc-060998abbc7a · outbound

This paper cites On Limitation of Transformer for Learning HMMs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On Limitation of Transformer for Learning HMMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.900610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.900610Z digest=sha256:674dfd21445588d6765b05a053c198033b45220f754a6c216225f116448a6c50

Observation 7a8e3379-4fcd-4c62-a39f-9cf93381c4c3 · outbound

This paper cites In-context convergence of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization In-context convergence of transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.856363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.012937Z digest=sha256:1494617c67d024cf18de5dde8912258c32e80c0cea7949cbedad606c09e5c86d

Observation 0f2c682a-00e9-44c6-927e-990e382e8067 · outbound

This paper cites How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.576290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.140528Z digest=sha256:3ed1ecddcdccd713629fe65e919b90ba5ffd803a09cdb4bcd3d6973da082dd31

Observation a492e44c-f911-481a-84ba-313a963d527c · outbound

This paper cites Vision Transformers provably learn spatial structure.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Vision Transformers provably learn spatial structure

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.476399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.267665Z digest=sha256:177cb13a1572fa06f2895c67bb39b7b612aa1c3c19fa7f09e8df0fc8855e6591

Observation 3bc12942-9cb9-4d32-8cc7-2ca58c640eff · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 26

Resolution
verified exact
doi, observed 2026-08-07T06:10:21.958846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.493685Z digest=sha256:c36643ee5a4668b96117f794640b8cbf33e52ecf78e2b75957b199aab469d76d

Observation daa10853-5dcd-416e-a272-2b2af1fbdbb7 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.579936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.579936Z digest=sha256:7dcb04029f2b5e6ef06ca800cbc116e5538497fde75bc756987ce3fa18e445e9

Observation f43d11c9-bdd9-46cc-ab9f-22c1271550ef · outbound

This paper cites PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.702105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.702105Z digest=sha256:cf9e403173b938509bc4a185a4c86f48e6535078ec6ea339147af60e1355b22a

Observation 1f9a66d0-e2b3-4bc2-8315-fd07ad268ec9 · outbound

This paper cites and Sato, I.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization and Sato, I

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.310452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.824938Z digest=sha256:31724d74de0d4fe1afd46ce20da09f5e21759e7a295290bd61f7c6cd120e3506

Observation 690e44ed-757d-411e-868e-ae61c42ee29c · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.920713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.920713Z digest=sha256:bde7e62ab7f4e170e19eb11e938a342e03a1c35152e1db811f9848b8d507ea07

Observation 6a9d517c-5ac8-4e1f-a088-3fbba6efc302 · outbound

This paper cites SimA: Simple Softmax-free Attention for Vision Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization SimA: Simple Softmax-free Attention for Vision Transformers

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.317828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.033907Z digest=sha256:9bbd822cfe91794a48818b45366ef0f661898c5120106e5c7724a6036cce1c40

Observation 4e9af55a-da66-4b5e-b0a8-190a9c3b0978 · outbound

This paper cites A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.105823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.105823Z digest=sha256:7466140f693c8cf44970ebd4266304921a7f6533ad585f3a0127f78e593944a0

Observation 10d5a435-75ec-4bac-ab9e-e3aad03ec028 · outbound

This paper cites The Closeness of In-Context Learning and Weight Shifting for Softmax Regression.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization The Closeness of In-Context Learning and Weight Shifting for Softmax Regression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.205176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.205176Z digest=sha256:e8ac2f9c2db97ad6976cb14ceaf89a0fdb9145b36121a0e456caf8c6400d6f77

Observation a19f856e-914a-4962-bedf-4c7694a05efa · outbound

This paper cites On the Expressive Power of Self-Attention Matrices.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Expressive Power of Self-Attention Matrices

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.335874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.335874Z digest=sha256:f223529168eeb395376bc72ecf7ef11f7a0c2fb6da71c7e052bf7e9a8c0694a8

Observation 4de1ab7e-9de2-452b-9150-05f154133b75 · outbound

This paper cites Transformers Learn Shortcuts to Automata.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Learn Shortcuts to Automata

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.422720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.422720Z digest=sha256:33419060bbdaa14a58c32ab203c1bd2e6ded2d8cf7439a95c5762a0d2854879b

Observation 4a00c977-f4d6-4b39-9d34-8f98f198fe77 · outbound

This paper cites Rethinking Transformers in Solving POMDPs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Transformers in Solving POMDPs

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.099785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.505098Z digest=sha256:26b216b60f26f0a0a8b23255fc71a1f66d42af3b62a6d32c57249a74a71c5e55

Observation 609770b0-6828-4366-8570-e6ee8d580761 · outbound

This paper cites Your transformer may not be as powerful as you expect.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Your transformer may not be as powerful as you expect

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.114111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.623244Z digest=sha256:49753d1b8f4cec6699a6faf916bb19fb196188cba128b0ca8a7e1737d358ba04

Observation eccb2b48-c5b2-4ff5-abaa-3473c20fd9ba · outbound

This paper cites Transformers are Expressive, But Are They Expressive Enough for Regression?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Expressive, But Are They Expressive Enough for Regression?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.708151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.708151Z digest=sha256:21686034283c03e84c528db0ed357aa72c32667e44c5d07d2f56284ea7debb76

Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · outbound

This paper cites Theory, Analysis, and Best Practices for Sigmoid Self-Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.778678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.778678Z digest=sha256:516bfb9979e2d4a6cd2d31dfc43994de43e2547fe270391d329442af168ac356

Observation 9e29a8a8-5a31-4110-b43b-ed1dbccdd87e · outbound

This paper cites Representational Strengths and Limitations of Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Representational Strengths and Limitations of Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.882192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.882192Z digest=sha256:b38cd604fe4f9d0a6ffb28cdfd820b4c20eef9b8f5c1252a5d4626767e0459ce

Observation a1ee5037-49bf-4eaa-844f-c146f8fa6315 · outbound

This paper cites Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:21.767656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.014640Z digest=sha256:e2b8442a515e2ebc4d550bced0ccb7cb71f71f79fb877fcbec8aa9ed81ae2788

Observation db8e1eb0-9a49-4aa4-8bee-0609d5641170 · outbound

This paper cites Unraveling the gradient descent dynamics of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unraveling the gradient descent dynamics of transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:23.887917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.096084Z digest=sha256:dc7ed07a39a7a7882f8631c6ac81f0bdbcfaf340c29e7fc357c7501d64d7924c

Observation 67c703c1-8a6b-4f82-be50-7ba0e106a85f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.168393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.168393Z digest=sha256:8ebe64dcdd852ac1b430253eef1027d546120a275e49a93c54929ddfac8c6f20

Observation 2b50d96b-e00c-4ba8-9ca7-4a4b3b64e704 · outbound

This paper cites Transformers as Support Vector Machines.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers as Support Vector Machines

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.263719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.263719Z digest=sha256:dd95018571823355a43cef677a356bfbcc06e7e76f19ec6a6900c90915fe3a69

Observation ec4bcdc5-53d4-47c7-a369-99fd34080d22 · outbound

This paper cites Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.436184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.436184Z digest=sha256:fb2ab4522804492ba9a5cf0348605b94d8febbc69b740ad5ff3fadc4931d1407

Observation c2bc851c-fdb9-4894-8e9b-b6dbb3bdb473 · outbound

This paper cites An Introduction to Matrix Concentration Inequalities.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Introduction to Matrix Concentration Inequalities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.545869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.545869Z digest=sha256:f1cf7bf142d250cfabcb9f780d3109752d38fba9205fc96bb2d820fdd767bca9

Observation 7d115a87-ac16-4df7-9565-0ad748ae8882 · outbound

This paper cites Attention Is All You Need.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Attention Is All You Need

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.677748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.677748Z digest=sha256:32046db7c2cbdfdd4866c0fb93ddc4403df4a44a47aa6c38fa83b57122b9662d

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · outbound

This paper cites Transformers learn in-context by gradient descent.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:a7701cfb514599ac56683b35dd31bb52ec476c2cbb1f89fdbe87f5071606d568

Observation ce0005a2-a080-41ca-ba0c-86ce5636c97f · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linformer: Self-Attention with Linear Complexity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.871924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.871924Z digest=sha256:29b21c2320d61a5a79a0802d52c75c59b995599463f4058dbb55677979d68d5f

Observation 30eb5080-6990-49a8-882f-2efd1e737368 · outbound

This paper cites Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.984325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.984325Z digest=sha256:8d37a360e3f9ef9cd0f773090b90f92258d9850cfa2eda01dffd3a262aae6a9b

Observation 4559da0e-1c67-4936-926e-00f186f34955 · outbound

This paper cites Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.092976Z digest=sha256:209a82628501f626153a77fe3d77aa99ff0be4c528554c1f9de0de917aa9c393

Observation 3c471eee-41ae-4717-a298-d15b0b9f958f · outbound

This paper cites Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.715533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.173804Z digest=sha256:eb0d7e685bb49c00a7afe25b52760c6812deb81ca3e8b42dc11694b8e33a6112

Observation 65e87cbc-8b73-4884-8d06-560981f90fff · outbound

This paper cites Self-Attention Networks Can Process Bounded Hierarchical Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Self-Attention Networks Can Process Bounded Hierarchical Languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.255804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.255804Z digest=sha256:6b312ab3c8e6b767e133c68307242262ed7212e189933ebcbabcbf0739aedf01

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:aaedc5d5d12d01f560b980ab5bf6fb742b8c6973eb64e90a74a1fbbacd468f26

Observation 811b1712-1d5b-43dd-8bef-2b8dfe8355f3 · outbound

This paper cites Global Convergence of Block Coordinate Descent in Deep Learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Global Convergence of Block Coordinate Descent in Deep Learning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T06:10:22.512246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.459603Z digest=sha256:41ce346b9e8b43556274a48f076c8775c0e8662df0bb79126e1ebd0081ac5294

Observation 1374c329-c8d2-4784-9261-e9b11af7be25 · outbound

This paper cites Transformers are Efficient Compilers, Provably.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Efficient Compilers, Provably

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.360182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.549139Z digest=sha256:12dc5027ba960e6ca58a4b73eb225438a57f46a11b569d6b50446676a2c0fa2e

Pith citing papers

No inbound Pith citation observations are available.