Pith. sign in

Paper Citation Record · LEDGER

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.06179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06179 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:21.549139Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact8
  • verified fuzzy7
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b186e2f3-97c6-49a0-869d-367bed3f96c2 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.045295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.045295Z digest=sha256:ac9117f7ead668c6397848fd9591225b1bad0d996529273ec4cb68eff381b4a4

Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.118380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.118380Z digest=sha256:325311e9db0b90b752f9dbd48b7c7ea568b736c3abd11489676b3389d3d55002

Observation b86348a3-4561-4195-b6ea-de5fbd1c5f02 · outbound

This paper cites Block coordinate descent for neural networks provably finds global minima.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Block coordinate descent for neural networks provably finds global minima

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.632599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.308487Z digest=sha256:5b09fb1b85d2c0e58e0b1d4be9c4565f261e6fe13615adef3324b72dfcba6382

Observation ed9c1163-000a-4186-bcff-5fc84f90fc4f · outbound

This paper cites How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.430307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.430307Z digest=sha256:4ef29b7a10fee79e797cfbb08832be5cd913a8fb288f8dbbc6777f807caac4c4

Observation e07cc56c-df49-44ff-8f9e-95ae7a63a9e7 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.520290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.520290Z digest=sha256:e7c7fbea8078640d47ed95711abd32e4708b37a256349a892c9186516bb71586

Observation 10cf59d9-00e8-4344-8bf5-a1783691c633 · outbound

This paper cites On the Ability and Limitations of Transformers to Recognize Formal Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Ability and Limitations of Transformers to Recognize Formal Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.604932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.604932Z digest=sha256:8af5150ba8a4f584f91fd0ea3194ec5d5765cd118d32eb43926f0ef99491d369

Observation 3da02f76-adfb-40e5-a345-516b08524857 · outbound

This paper cites On the Computational Power of Transformers and its Implications in Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Computational Power of Transformers and its Implications in Sequence Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.691808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.691808Z digest=sha256:e106549d9535bbf3ebde9bdcbed05f3ec14516c741c7d7435fffe0df48b97e88

Observation 7e95b74c-6263-4093-bc2a-4b82c380afc1 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:10:25.347936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.756489Z digest=sha256:2dd5c0cf09122da12fbc1357e532cd92e2d52037cc5a37f7fa04c9e096e5f2ed

Observation d1e61cd4-1248-4fe4-9c20-fc4aaaf28195 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.825748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.825748Z digest=sha256:ce3fe5256035f40f0e1460482c82b8959f58f442edb39095bfcc89b189ab9ae1

Observation 536d2bdd-a47a-4ba5-935b-198c6aa00d06 · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.886542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.886542Z digest=sha256:d0918fd47f3d159b8696a38640bd32f319dc1516930aedd5f0cd4924b7005f56

Observation 56b1e1c8-31ae-4eb2-a06f-9c493c7e724e · outbound

This paper cites Provably learning a multi-head attention layer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Provably learning a multi-head attention layer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.973138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.973138Z digest=sha256:8396dbfcb610b811680f9a71015db38f94243c86d9aef1283e28a88b5af46f10

Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.068890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.068890Z digest=sha256:25c0dde1207e9ddb64776638edeec53f394e7a8dfd8325a79cfcdbb7370d791d

Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · outbound

This paper cites Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.138258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.138258Z digest=sha256:da13bde78a49310f88c106a1e95386b8dd16033fa728d9c6135acb7e1a839e3e

Observation 6a550f40-a8a8-4ee9-b9cd-7b0e0e2fd991 · outbound

This paper cites Rethinking Attention with Performers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Attention with Performers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.231632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.231632Z digest=sha256:cf30b7767f880bd1eb3cd15bbd2592b0729ed59f6c02bd2cf39cfa6c466f54f9

Observation cd45c953-249f-42de-9f57-f256a71e4bd0 · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Optimization and Generalization of Multi-head Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.345237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.345237Z digest=sha256:dd48d0b40da1e10ea6a266b878887dbe0fa2711ef92e4bbef5354fab425ac641

Observation d9661c3a-1b2a-4d0c-a859-bb38637c1eae · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.408915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.408915Z digest=sha256:9c35d093b484ab88e24632c34987b3269a9de7387d3c666f2ceff4aeed22c944

Observation f99701b9-2adf-497f-9573-589492dc5b38 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.511563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.511563Z digest=sha256:26898c94fbcfc2a41c9c08b5b3b49023e40a914d2383b922e5481de060ef6d06

Observation 4d4f7717-7ac1-4bc6-b735-738760779656 · outbound

This paper cites Inductive Biases and Variable Creation in Self-Attention Mechanisms.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Inductive Biases and Variable Creation in Self-Attention Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.582759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.582759Z digest=sha256:70cea04a07d67435cdf84a842e6baba8637a61f384a9fc680e6e2c2a31117aff

Observation 88ce5238-6b0a-4348-94a3-7f4802a095c6 · outbound

This paper cites A mathematical framework for transformer circuits.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A mathematical framework for transformer circuits

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.672307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.672307Z digest=sha256:d1d3802e68d93ac8c27498b62a0cdccf6ccddc445c6aed124b4c2c9fc001cb7c

Observation 31ae237e-f46c-40aa-868c-c80d2b355632 · outbound

This paper cites Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T06:10:22.175983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.767563Z digest=sha256:fb7e557d4c2ae292b407b335b054eafc52e7634f2f3ab59a7b4fab279f1b1bbf

Observation 22bbf9eb-1c4c-4bc6-92ea-a01c9e52edb3 · outbound

This paper cites M., and Fan, J.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization M., and Fan, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.117893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.829721Z digest=sha256:e460d3e246f0ee55172a1ae29b623b69eba169aa04d89b0868f8417acc23cdeb

Observation 6362e547-8525-4924-bdcc-060998abbc7a · outbound

This paper cites On Limitation of Transformer for Learning HMMs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On Limitation of Transformer for Learning HMMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.900610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.900610Z digest=sha256:2fc2098afbccfbfe95e335f290aebaa9a01214cb890b5ed1e649a1d3c320c9ca

Observation 7a8e3379-4fcd-4c62-a39f-9cf93381c4c3 · outbound

This paper cites In-context convergence of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization In-context convergence of transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.856363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.012937Z digest=sha256:f1acfeff9147bf7a94363034cde008fe98a6d5e47ea518e9016d1ed92eb3fbce

Observation 0f2c682a-00e9-44c6-927e-990e382e8067 · outbound

This paper cites How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.576290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.140528Z digest=sha256:b1a11ab6695f42146bb10b97c5f592900643481d09fb6f53f6ae4432003cd5ab

Observation a492e44c-f911-481a-84ba-313a963d527c · outbound

This paper cites Vision Transformers provably learn spatial structure.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Vision Transformers provably learn spatial structure

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.476399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.267665Z digest=sha256:5519ced096e1923d512fa122596222d2ed7924d8c63ce5afb793dc4f71b8e670

Observation 3bc12942-9cb9-4d32-8cc7-2ca58c640eff · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 26

Resolution
verified exact
doi, observed 2026-08-07T06:10:21.958846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.493685Z digest=sha256:348154d2f3bae13b47f06cd28bdfbb768c4bef14ac385405e3f8bea7733fa043

Observation daa10853-5dcd-416e-a272-2b2af1fbdbb7 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.579936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.579936Z digest=sha256:ec8b4d0bb8944bac85dd39bd45a2e77e563af2488d28a64823967d6b2543657a

Observation f43d11c9-bdd9-46cc-ab9f-22c1271550ef · outbound

This paper cites PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.702105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.702105Z digest=sha256:dec2ad80e63fc67154d3f6da680d385f244b78bbc1c15619980344a962fa2af9

Observation 1f9a66d0-e2b3-4bc2-8315-fd07ad268ec9 · outbound

This paper cites and Sato, I.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization and Sato, I

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.310452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.824938Z digest=sha256:c4295c76c0c69f5c36a55309171c9793089ec66c122cf82b979ce7445654d62b

Observation 690e44ed-757d-411e-868e-ae61c42ee29c · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.920713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.920713Z digest=sha256:512c4ee2e77069867a9dfe50fe609673c905811a2be7c3379a0a91e0bba6ed54

Observation 6a9d517c-5ac8-4e1f-a088-3fbba6efc302 · outbound

This paper cites SimA: Simple Softmax-free Attention for Vision Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization SimA: Simple Softmax-free Attention for Vision Transformers

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.317828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.033907Z digest=sha256:89b5ba1e600e0e342ec951f3bbe8922802ba02cebca8ce3b8b7aea62e4df6b0a

Observation 4e9af55a-da66-4b5e-b0a8-190a9c3b0978 · outbound

This paper cites A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.105823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.105823Z digest=sha256:77835dfdc139a5174a0ac57d80813675f8c7000751cb0dddef48eb63676eeb4f

Observation 10d5a435-75ec-4bac-ab9e-e3aad03ec028 · outbound

This paper cites The Closeness of In-Context Learning and Weight Shifting for Softmax Regression.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization The Closeness of In-Context Learning and Weight Shifting for Softmax Regression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.205176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.205176Z digest=sha256:f8cd212b77865216c31ce4014c2becd207fc2d218c2276abac80878411651eb5

Observation a19f856e-914a-4962-bedf-4c7694a05efa · outbound

This paper cites On the Expressive Power of Self-Attention Matrices.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Expressive Power of Self-Attention Matrices

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.335874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.335874Z digest=sha256:7c5b7d91d619bf40aa2e56517514a0e1b5e318de7b87f9d9891a141001a1ad80

Observation 4de1ab7e-9de2-452b-9150-05f154133b75 · outbound

This paper cites Transformers Learn Shortcuts to Automata.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Learn Shortcuts to Automata

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.422720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.422720Z digest=sha256:a9ff5dd0c1084cd8d2754cdaf2d3fa4bca43eb6fd13550c69004767ee5fae800

Observation 4a00c977-f4d6-4b39-9d34-8f98f198fe77 · outbound

This paper cites Rethinking Transformers in Solving POMDPs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Transformers in Solving POMDPs

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.099785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.505098Z digest=sha256:557de8b1c4a3457fd60406b5c8bf14471f0006dc541cbce010ff2aba26f3d38f

Observation 609770b0-6828-4366-8570-e6ee8d580761 · outbound

This paper cites Your transformer may not be as powerful as you expect.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Your transformer may not be as powerful as you expect

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.114111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.623244Z digest=sha256:9021cc612a3fef9a6eca954e051f130b1fb91495295eb540f9cf11fc7db087b0

Observation eccb2b48-c5b2-4ff5-abaa-3473c20fd9ba · outbound

This paper cites Transformers are Expressive, But Are They Expressive Enough for Regression?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Expressive, But Are They Expressive Enough for Regression?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.708151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.708151Z digest=sha256:bba59bac2bae8dad3478152743edb6c9702c34777aabe42c3728d6a6778152c2

Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · outbound

This paper cites Theory, Analysis, and Best Practices for Sigmoid Self-Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.778678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.778678Z digest=sha256:a9914fd38a4a43179ded0716840a42129ca7cd2c16112ee213c0bf53a7bca420

Observation 9e29a8a8-5a31-4110-b43b-ed1dbccdd87e · outbound

This paper cites Representational Strengths and Limitations of Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Representational Strengths and Limitations of Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.882192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.882192Z digest=sha256:54ec3514f9d1d1bbed838934df5c3e091bc9e30e82d13326ce4d6c5931d08e5e

Observation a1ee5037-49bf-4eaa-844f-c146f8fa6315 · outbound

This paper cites Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:21.767656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.014640Z digest=sha256:bc82b9fa4aea8bc1e60cfab83237a53e2bc48713568857065e2b2edd0ff8b6d0

Observation db8e1eb0-9a49-4aa4-8bee-0609d5641170 · outbound

This paper cites Unraveling the gradient descent dynamics of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unraveling the gradient descent dynamics of transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:23.887917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.096084Z digest=sha256:c67762a837abe3cca3e486f2e0c687d38f65191feea2b7492253d753c079ea73

Observation 67c703c1-8a6b-4f82-be50-7ba0e106a85f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.168393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.168393Z digest=sha256:7ae3f6edd60c8dffd228893972a323eb1babf28114f108526601709d5b68f419

Observation 2b50d96b-e00c-4ba8-9ca7-4a4b3b64e704 · outbound

This paper cites Transformers as Support Vector Machines.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers as Support Vector Machines

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.263719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.263719Z digest=sha256:3ffaacdbc1b0fc0403650966c40b4fa18a2bacbace9eb0bb492189534dda9f78

Observation ec4bcdc5-53d4-47c7-a369-99fd34080d22 · outbound

This paper cites Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.436184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.436184Z digest=sha256:a2906cbb23fb6ec593c344f1d896df0dd17dbc50954e361031fba18ea4be14cc

Observation c2bc851c-fdb9-4894-8e9b-b6dbb3bdb473 · outbound

This paper cites An Introduction to Matrix Concentration Inequalities.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Introduction to Matrix Concentration Inequalities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.545869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.545869Z digest=sha256:1f768ded9dd944c8a803e837c5b807cb4a8962027644bd793c60e5812cce510c

Observation 7d115a87-ac16-4df7-9565-0ad748ae8882 · outbound

This paper cites Attention Is All You Need.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Attention Is All You Need

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.677748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.677748Z digest=sha256:87e5595f819350f8787efd70f44cffde1659c8967580eca23af7b3c5ce45516a

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · outbound

This paper cites Transformers learn in-context by gradient descent.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:075efa909faeea6df0ff6a38af75aa628bd1757ff4409f5eddd5f90c79a80a56

Observation ce0005a2-a080-41ca-ba0c-86ce5636c97f · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linformer: Self-Attention with Linear Complexity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.871924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.871924Z digest=sha256:b15bb7e192d0f4c07c8c5346c287f6359603f24c8a48f918ba0f641c1834ed6c

Observation 30eb5080-6990-49a8-882f-2efd1e737368 · outbound

This paper cites Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.984325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.984325Z digest=sha256:21f4a1a1ea514ad1e7b5bf0df87cf3245a78151084fa8ddd08592ca1516dfdb9

Observation 4559da0e-1c67-4936-926e-00f186f34955 · outbound

This paper cites Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.092976Z digest=sha256:07883ff9148c62aff40f8fca9ea29f092352af3e65a98c8a675810e189b25d71

Observation 3c471eee-41ae-4717-a298-d15b0b9f958f · outbound

This paper cites Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.715533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.173804Z digest=sha256:85d8ef319ab7d25c217497490824a34200db4d115fe34257c439b0a7f2c67032

Observation 65e87cbc-8b73-4884-8d06-560981f90fff · outbound

This paper cites Self-Attention Networks Can Process Bounded Hierarchical Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Self-Attention Networks Can Process Bounded Hierarchical Languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.255804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.255804Z digest=sha256:e96e2bda1416aec03668f84e359e1a7829cabbc62fbdbdcc7b6ae6faa44f2cbb

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:ed44eda78bbea3de002aad41cf498044e21e105a73fe10b8084daa57f3e482e7

Observation 811b1712-1d5b-43dd-8bef-2b8dfe8355f3 · outbound

This paper cites Global Convergence of Block Coordinate Descent in Deep Learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Global Convergence of Block Coordinate Descent in Deep Learning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T06:10:22.512246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.459603Z digest=sha256:60f303039ab0af933fd2d1c7e22655cfb5be6f0682b112136a28e9045fc32f43

Observation 1374c329-c8d2-4784-9261-e9b11af7be25 · outbound

This paper cites Transformers are Efficient Compilers, Provably.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Efficient Compilers, Provably

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.360182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.549139Z digest=sha256:f57f5c75a4aa8731823f7188bfc81e9c4d71d411f91d3b259a850dc4635f3b3e

Pith citing papers

No inbound Pith citation observations are available.