Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics of In-Context Learning in Linear Attention

As of 11 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 12 inbound Pith citation observations for arXiv:2501.16265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16265 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:42:09.164276Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:54:19.125912Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 863a75de-3fea-426f-beb9-9949540cbda1 · outbound

This paper cites write newline.

Training Dynamics of In-Context Learning in Linear Attention write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.771564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.771564Z digest=sha256:b344b7934ca601776babf0a59433522caca8d218c87d73c9d2f7dcf56916db8a

Observation 72351a8a-da9b-4801-a986-b92766eb20b9 · outbound

This paper cites Context-Scaling versus Task-Scaling in In-Context Learning.

Training Dynamics of In-Context Learning in Linear Attention Context-Scaling versus Task-Scaling in In-Context Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-10T13:42:09.943815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.777812Z digest=sha256:d65c0600a56892eb422fca530b37146c9eaddfbfffbaa248b48bba3f62d5be0c

Observation 64959c77-d4e0-4ead-8c55-ed7fb68768ca · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.784728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.784728Z digest=sha256:d6e722ae39a61aba7b6d3335941ac24c4640eb25b547038202dec457cc5e1f60

Observation 0e7d4324-81d1-4fc7-8515-0dcccf45abe0 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Training Dynamics of In-Context Learning in Linear Attention Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.789849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.789849Z digest=sha256:9e65473c1e1c3e41b8f38df129bcd8225dbc2091a8d79f9107484c1d030cf105

Observation e7b2f64c-7709-4c3f-b47e-b65656303604 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

Training Dynamics of In-Context Learning in Linear Attention What learning algorithm is in-context learning? investigations with linear models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.794517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.794517Z digest=sha256:10966aee350ab4da09a00f10e9f46b2cf96f74fb326b1c6ccdc1e41a5f87998e

Observation d12b03fd-e8ef-410e-9551-b1ffaaaf4593 · outbound

This paper cites A., Merullo, J., and Pavlick, E.

Training Dynamics of In-Context Learning in Linear Attention A., Merullo, J., and Pavlick, E

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.799444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.799444Z digest=sha256:7995faa105fcaad0f2efa3cc6a9ca2669cd0dcd73ccb6acd80efec2e77825132

Observation 5d7f1b81-2cdc-4b17-a1af-749a4644f8a3 · outbound

This paper cites Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens.

Training Dynamics of In-Context Learning in Linear Attention Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.805009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.805009Z digest=sha256:eb8ee98e0503884d2c2676d57d7cfea952e264f246a09723e64b809011e42447

Observation f7074a8a-00f3-4479-9263-5718ab8df39f · outbound

This paper cites A convergence analysis of gradient descent for deep linear neural networks.

Training Dynamics of In-Context Learning in Linear Attention A convergence analysis of gradient descent for deep linear neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.852752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.811298Z digest=sha256:5e4b6842cdb37d3fd1f543f32de9e78c8ef64c453589712cbe2c68cd9f2cc069

Observation 5c3b3432-1ef3-4765-aacf-9aef3aad3703 · outbound

This paper cites Max-margin token selection in attention mechanism.

Training Dynamics of In-Context Learning in Linear Attention Max-margin token selection in attention mechanism

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.838845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.815984Z digest=sha256:cfa9ad3364951c8e5141abdd4f8120b1b2dbfc6bfad39663718d7b92457c96fa

Observation 9cf46489-6ab3-41b9-bc27-ae77d20a4f8a · outbound

This paper cites Neural networks as kernel learners: The silent alignment effect.

Training Dynamics of In-Context Learning in Linear Attention Neural networks as kernel learners: The silent alignment effect

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.820551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.820551Z digest=sha256:dbb00bdd7e7b28e31c45f91a5a162cff0e2b15543f011a4d49806d9a0bf972ae

Observation 349988df-dfb1-4152-9151-a32999870656 · outbound

This paper cites Transformers as statisticians: Provable in-context learning with in-context algorithm selection.

Training Dynamics of In-Context Learning in Linear Attention Transformers as statisticians: Provable in-context learning with in-context algorithm selection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.815157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.825679Z digest=sha256:d20125115f776132945440cc70ea8f8dc65cdaf225c9b19399ea535885171c62

Observation bb2831a1-992f-4bca-9856-b729fbcd9d85 · outbound

This paper cites Transformers learn through gradual rank increase.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn through gradual rank increase

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.800154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.830426Z digest=sha256:ccaebd92733d1e4fa8318341d7e02742d69e89dd964ed1b4542bd430121b758b

Observation 66622513-0ceb-48fe-b15f-ce4584ee3eba · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.784524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.835092Z digest=sha256:00da596bdec8207462072ddc82fde2751357ea2b839739ec59bb9f71d6e93b1d

Observation a49640b7-93d8-485e-a725-2fe526ba29fd · outbound

This paper cites Infinite limits of multi-head transformer dynamics.

Training Dynamics of In-Context Learning in Linear Attention Infinite limits of multi-head transformer dynamics

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.770121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.839610Z digest=sha256:90efbb37cf11cdc4f5ebc4c509ce3cabb57aeff673bb2b0191563029dde4427b

Observation 4b316124-1a7f-4e01-a6df-cee59b3afa68 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.844836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.844836Z digest=sha256:8ab30f5e6ae1b5c6c2c2f0849db71286b38638ff9f21ffd8eccae4b9622f0c47

Observation 9431fafa-7de8-4929-b83c-34a635225f91 · outbound

This paper cites Toward understanding in-context vs.

Training Dynamics of In-Context Learning in Linear Attention Toward understanding in-context vs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.745978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.849357Z digest=sha256:3270384204ecef92481bbcc3784a682ade9855f2e9e733147739610909ba4417

Observation deec83fe-93d7-433b-9baf-a99c3158bce9 · outbound

This paper cites Data distributional properties drive emergent in-context learning in transformers.

Training Dynamics of In-Context Learning in Linear Attention Data distributional properties drive emergent in-context learning in transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.731984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.854192Z digest=sha256:a0acb26051df612df2de62109b28d109cc2ac9f858b4eb81391de1a935dc717f

Observation cdd4c5b0-3ed7-421d-bff1-ef95b33764c2 · outbound

This paper cites L., and Saphra, N.

Training Dynamics of In-Context Learning in Linear Attention L., and Saphra, N

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.718066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.858728Z digest=sha256:1572684c10e87bc632bde71200201c148cb002e95a2f7851fe8c8d122278ee4c

Observation 51982e89-d015-491b-8279-7a1ff550281b · outbound

This paper cites Provably learning a multi-head attention layer.

Training Dynamics of In-Context Learning in Linear Attention Provably learning a multi-head attention layer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.863656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.863656Z digest=sha256:d696bea4c4170da37565655a00955112533c4649d4dd3d322cda64dcfbc1fcd4

Observation 48c82c41-46f1-484f-9f4d-d176ffbd6535 · outbound

This paper cites Training dynamics of multi-head softmax attention for in-context learning: Emergence, convergence, and optimality.

Training Dynamics of In-Context Learning in Linear Attention Training dynamics of multi-head softmax attention for in-context learning: Emergence, convergence, and optimality

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.704119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.868864Z digest=sha256:d18528b1cf7412b55e1134a3923e8f1e63549178ea5918c86ed311e53719046a

Observation 2dfe84c6-4a23-4767-90aa-9e8e1a1b41e7 · outbound

This paper cites Unveiling induction heads: Provable training dynamics and feature learning in transformers.

Training Dynamics of In-Context Learning in Linear Attention Unveiling induction heads: Provable training dynamics and feature learning in transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.689327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.873299Z digest=sha256:9c145584b05b64052d6a1499c4aa76a741384c39f6b75aee58bd140d845d7769

Observation 8e394567-653e-4201-9626-cbb173be8e5a · outbound

This paper cites On lazy training in differentiable programming.

Training Dynamics of In-Context Learning in Linear Attention On lazy training in differentiable programming

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.674486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.877770Z digest=sha256:f57594754bd48c8e3077f4bd55f4f2440501299cc9d14c6d1b64347f39c09854

Observation e00d3fd6-135a-4b81-932c-f81f3fd0b8bb · outbound

This paper cites S., Hu, W., and Lee, J.

Training Dynamics of In-Context Learning in Linear Attention S., Hu, W., and Lee, J

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.659493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.882860Z digest=sha256:109b72310a142673224f6e8455f0b77d60e5f2842ad715b65af7e7ca876e888d

Observation eb1251fe-4819-42d0-b142-b4df18a1c9d9 · outbound

This paper cites Finite Sample Analysis and Bounds of Generalization Error of Gradient Descent in In-Context Linear Regression.

Training Dynamics of In-Context Learning in Linear Attention Finite Sample Analysis and Bounds of Generalization Error of Gradient Descent in In-Context Linear Regression

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.887996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.887996Z digest=sha256:6086a6fec621c36ef3795a82b8fd65fe113e4299ebafe93a5007d56716bc2490

Observation e1295f8f-2108-4b23-8d35-5584f6f90474 · outbound

This paper cites The evolution of statistical induction heads: In-context learning markov chains.

Training Dynamics of In-Context Learning in Linear Attention The evolution of statistical induction heads: In-context learning markov chains

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.646077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.893622Z digest=sha256:0157f63b4e0855dd66a984812c645fee587bd8cc74608564f35b732c674d0107

Observation 21e847e9-0a87-4176-813d-e0ee1a1e4fb8 · outbound

This paper cites and Vardi, G.

Training Dynamics of In-Context Learning in Linear Attention and Vardi, G

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.630905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.898470Z digest=sha256:9746ed41e918b0e4fff027948fa9d74faee256d3e5b2f94d97bfedf100264d16

Observation 71849b5f-4342-4594-9571-5e9f60e61493 · outbound

This paper cites Transformers learn to achieve second-order convergence rates for in-context linear regression.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn to achieve second-order convergence rates for in-context linear regression

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.617623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.903611Z digest=sha256:279e6544a51c9c29ab3752cb4815345c6868c73b2d062e09279dd19d46458176

Observation 5fce92da-7176-4c50-a44b-b8e47b0f7504 · outbound

This paper cites Effect of batch learning in multilayer neural networks.

Training Dynamics of In-Context Learning in Linear Attention Effect of batch learning in multilayer neural networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.602828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.908134Z digest=sha256:8d0cb79772059583a4d5fe7db618e24390f3013c1de2497c1f29150180430dff

Observation d3717c82-e22d-4d4a-a061-78241bdef035 · outbound

This paper cites S., and Valiant, G.

Training Dynamics of In-Context Learning in Linear Attention S., and Valiant, G

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.587295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.912602Z digest=sha256:1e0bfbee8d5c5a981d26b83140722c400a8edaaa114ef39ff140c01a0cf6e9a7

Observation 22b504e4-4ff0-43fe-b746-85863551cc3e · outbound

This paper cites On the Role of Depth and Looping for In-Context Learning with Task Diversity.

Training Dynamics of In-Context Learning in Linear Attention On the Role of Depth and Looping for In-Context Learning with Task Diversity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.917150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.917150Z digest=sha256:254161b498e6cca2ce286031803bd1021eca67ca29cfb8febfeddd1e583f1788

Observation 080117ce-e8b2-4bcb-b8ca-b821c4b05ab7 · outbound

This paper cites Dynamic metastability in the self-attention model.

Training Dynamics of In-Context Learning in Linear Attention Dynamic metastability in the self-attention model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.922647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.922647Z digest=sha256:8b75c8d8120e1c611acf79a879055751811156a5aeaade44ddbf198317354e54

Observation a1fc2764-3e10-4b0f-8f81-e906ccb50c9d · outbound

This paper cites In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention.

Training Dynamics of In-Context Learning in Linear Attention In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.927759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.927759Z digest=sha256:b98ed675b9dc89128a10ff1aeeac4fb51ec7624f53273127e8074d7decfee6fb

Observation 6860ba27-79d8-4f9a-b077-ae624ebe29f6 · outbound

This paper cites Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks.

Training Dynamics of In-Context Learning in Linear Attention Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.572786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.934151Z digest=sha256:a69c6cc5a93be939300c9b966379d5978d1777d65c558c9c241186a63b4f27c9

Observation 0e993c53-abc7-4b3a-84a4-c351153cc72a · outbound

This paper cites T., Schrodi, S., Bratuli\' c , J., Behrmann, N., Fischer, V., and Brox, T.

Training Dynamics of In-Context Learning in Linear Attention T., Schrodi, S., Bratuli\' c , J., Behrmann, N., Fischer, V., and Brox, T

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.557977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.938582Z digest=sha256:34778bb16c63140840012bb227e7f87d27da01a8098fa53508e1b9f4eb8f75c0

Observation acf3c34c-7f88-4ad2-98a2-954ff45b44d5 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.543221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.943036Z digest=sha256:2d8f1cfe0b669c6590a80ba53975c3e2b9367bbedaa4ccd46e727b0981e8f770

Observation 7ecd2b09-c6dc-49c3-ba91-125b688ed482 · outbound

This paper cites Non-asymptotic convergence of training transformers for next-token prediction.

Training Dynamics of In-Context Learning in Linear Attention Non-asymptotic convergence of training transformers for next-token prediction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.528891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.947058Z digest=sha256:55ba8d7aa7bf026cac75561146fb7de0c913071daf380fe3955463d56848dd2e

Observation f74697cf-cb95-4040-afec-0b0370f1f431 · outbound

This paper cites In-context convergence of transformers.

Training Dynamics of In-Context Learning in Linear Attention In-context convergence of transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.513722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.951456Z digest=sha256:ecf0d95ee6aace3838a846414e955c975ec30c56b4ccc7534664267f5e8802de

Observation 7d00bd3f-797c-4623-93a3-701fea851152 · outbound

This paper cites A theoretical analysis of self-supervised learning for vision transformers.

Training Dynamics of In-Context Learning in Linear Attention A theoretical analysis of self-supervised learning for vision transformers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.499288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.956251Z digest=sha256:2af7ced7696085e056266f38f969689e562346f9b702aabb9291f5a908338c7f

Observation a1af68c1-b9e9-4290-a60a-198e97f63f17 · outbound

This paper cites E., Huang, Y., Li, Y., Rawat, A.

Training Dynamics of In-Context Learning in Linear Attention E., Huang, Y., Li, Y., Rawat, A

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.483227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.960781Z digest=sha256:617040eec3d780d84c1f8a2964b65f957d75112d21b91e5a17420b25d922afd3

Observation 94160fa6-dc10-40e2-96d4-fee0712934d6 · outbound

This paper cites D., and Ryu, E.

Training Dynamics of In-Context Learning in Linear Attention D., and Ryu, E

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.462332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.965195Z digest=sha256:d321a32593e774d04e3adc0c920e873ec71004b3feaf28c07f1748d17fec2d5f

Observation d8050a11-f76c-47c8-a94a-6973553b1eec · outbound

This paper cites Vision transformers provably learn spatial structure.

Training Dynamics of In-Context Learning in Linear Attention Vision transformers provably learn spatial structure

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.446003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.969754Z digest=sha256:cdaec8b061969afc068773a7f14b3ff31bb530ba7b6212c7828176f36835c642

Observation d3c3ecfa-dfca-41db-89e8-8b638a718ee0 · outbound

This paper cites and Telgarsky, M.

Training Dynamics of In-Context Learning in Linear Attention and Telgarsky, M

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.430027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.974419Z digest=sha256:17659bab262ae768aab1b5531c7d198f27e9ebed63b61fe23401af5d54e3459a

Observation a5dff65c-27b7-4a06-8c01-740afd2959a5 · outbound

This paper cites Unveil benign overfitting for transformer in vision: Training dynamics, convergence, and generalization.

Training Dynamics of In-Context Learning in Linear Attention Unveil benign overfitting for transformer in vision: Training dynamics, convergence, and generalization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.412050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.978996Z digest=sha256:9fc86a8cf23f7aa6deca737f00f718057572ca98a72bd07f81fbba8bf6500ff3

Observation 91ed1275-8bee-458e-8065-2b87698e280a · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.983616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.983616Z digest=sha256:1fc2dac7afb9439cfc4c4c9a7397b8d7c49d8dc30338f0b8bead6132b9a8a1b8

Observation 050d0fdb-a8db-43ce-a50f-a50a63825831 · outbound

This paper cites and Suzuki, T.

Training Dynamics of In-Context Learning in Linear Attention and Suzuki, T

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.394142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.988243Z digest=sha256:93c79cdef49327faa3502df878d5da7d10a2d5f1c2f088bb17157bdb706153dc

Observation 5b6df4c7-96ab-467b-b6ef-96bda7b3e719 · outbound

This paper cites Geometry of linear convolutional networks.

Training Dynamics of In-Context Learning in Linear Attention Geometry of linear convolutional networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.992917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.992917Z digest=sha256:b558b953cb2677f9919109b8aca23218a6e1a11ec4b2d45335e10f58dadf5a0a

Observation 7adbada2-0e9e-4892-9d7d-97f8fc51d6cd · outbound

This paper cites Function space and critical points of linear convolutional networks.

Training Dynamics of In-Context Learning in Linear Attention Function space and critical points of linear convolutional networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.997544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.997544Z digest=sha256:389f4da33ecd38c44a700982e84b95c8c9f48f6356f61e3bc5a5a9dc21933e48

Observation ac224662-0663-4de0-ad94-82ec03472059 · outbound

This paper cites Is attention required for ICL ? exploring the relationship between model architecture and in-context learning ability.

Training Dynamics of In-Context Learning in Linear Attention Is attention required for ICL ? exploring the relationship between model architecture and in-context learning ability

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.377801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.002949Z digest=sha256:6cd59aa83ffbe19cd19dfd7295d5dbc228816f45a697c0f255327b16be8bb88e

Observation 1c9dc810-147c-4185-8fc3-e45f9cee7c70 · outbound

This paper cites Fine-grained analysis of in-context linear estimation: Data, architecture, and beyond.

Training Dynamics of In-Context Learning in Linear Attention Fine-grained analysis of in-context linear estimation: Data, architecture, and beyond

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.361926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.007960Z digest=sha256:67935e66ee4ade6e695cfe90bc76eda3bb31571a3dafc3a95dbe315b9e793f30

Observation bfcc9474-cd07-44c3-a664-9f1530f7c447 · outbound

This paper cites M., Letey, M.

Training Dynamics of In-Context Learning in Linear Attention M., Letey, M

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.012085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.012085Z digest=sha256:775462e03897c914cb0b6cc70034287568c541c4bf66ddee3e0f46d56fb606d1

Observation fb17ad90-e3da-42b1-9f02-1f4e87fa2969 · outbound

This paper cites V., Hashimoto, T., and Ma, T.

Training Dynamics of In-Context Learning in Linear Attention V., Hashimoto, T., and Ma, T

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.345757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.016999Z digest=sha256:4d2db3afaaa154a368b7633d7ca616169b1667e52d2ffd7b017a77fb8fc71967

Observation 35bb2cd7-dc34-417c-ad0d-6a6a4981f5cc · outbound

This paper cites V., Bondaschi, M., Girish, A., Nagle, A., Kim, H., Gastpar, M., and Ekbote, C.

Training Dynamics of In-Context Learning in Linear Attention V., Bondaschi, M., Girish, A., Nagle, A., Kim, H., Gastpar, M., and Ekbote, C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.328119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.022056Z digest=sha256:80bafa5f3599dbf06b5e5f64b1414554a579ab1d8103d6b5d81dd77c3989480a

Observation 5b4a75ca-26e7-4ecb-9977-4d41283da5f6 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Training Dynamics of In-Context Learning in Linear Attention Progress measures for grokking via mechanistic interpretability

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.026844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.026844Z digest=sha256:3611b2ff8d9dc4cd9c99c7a1a6a2edb5f23985f165e09a229649b7fb2fc71e9f

Observation daa75e13-4970-47e2-8c0e-02553d455f70 · outbound

This paper cites and Reddy, G.

Training Dynamics of In-Context Learning in Linear Attention and Reddy, G

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.299913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.031612Z digest=sha256:a2cea475e17beb50788e1baa8b254c43ef571a15de7e5f62e618b62e2a9c3a10

Observation 927c95b6-68bb-4d4e-91d3-68fda00b5b24 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.280694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.036312Z digest=sha256:eb5abd48872be59b268c041fbeb985117c09fbb45d00defadc07482443b37420

Observation 98433b6e-25fe-4c04-a3e6-b8278de044cb · outbound

This paper cites In-context learning and induction heads, 2022.

Training Dynamics of In-Context Learning in Linear Attention In-context learning and induction heads, 2022

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.263268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.040604Z digest=sha256:8bd52838cc2905e2040ddd9b6260c5e78c7eb430558426fa3eaebf5307bead2e

Observation 161f5a43-37d3-47d4-bb9c-813d4208ea58 · outbound

This paper cites and Reznikoff, M.

Training Dynamics of In-Context Learning in Linear Attention and Reznikoff, M

Reference 57

Resolution
verified exact
doi, observed 2026-08-10T13:42:09.225530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.044925Z digest=sha256:24a55bece03f85fcaba6893985d8b1dc45720e183e749c82b45b765a29e086a6

Observation ef1435cf-fe21-443d-a782-3920db0624bc · outbound

This paper cites F., Lubana, E.

Training Dynamics of In-Context Learning in Linear Attention F., Lubana, E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.240533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.049180Z digest=sha256:c24f36b0ee06ad40cb12a0b5c4b6f7d5caa1a1ac651aa0ca9ab77277b05b44e5

Observation e69add26-a54c-4ff5-98e6-f969c065b3df · outbound

This paper cites The mechanistic basis of data dependence and abrupt learning in an in-context classification task.

Training Dynamics of In-Context Learning in Linear Attention The mechanistic basis of data dependence and abrupt learning in an in-context classification task

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.053680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.053680Z digest=sha256:e67078a952b7c89f0a534cff25253040ee95d09348fec04a0edf2702de4939ec

Observation f49ea934-b4d3-490d-a612-247e02399420 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.213412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.058068Z digest=sha256:092ab80cae6e21efe6c08d7c930417ceaea505a91e7e72e86928e0d3cf036154

Observation 2d75092c-e492-4936-8c93-c2574e5815d2 · outbound

This paper cites A distributional simplicity bias in the learning dynamics of transformers.

Training Dynamics of In-Context Learning in Linear Attention A distributional simplicity bias in the learning dynamics of transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.199875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.063049Z digest=sha256:0d29f275fbd428e4ea7f49d23469e3d1d2914197fbee5fea5c9c27fc3368998d

Observation 33e03add-e130-498c-aabe-26e79c5e8aa7 · outbound

This paper cites M., McClelland, J.

Training Dynamics of In-Context Learning in Linear Attention M., McClelland, J

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.186428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.067331Z digest=sha256:dd632f3421b96203eeece66950833c8eb37c00ee3c9230fa9509b4b70c063cfa

Observation 0ae0ce06-93b8-44f8-85f5-afdeba87eb8f · outbound

This paper cites M., McClelland, J.

Training Dynamics of In-Context Learning in Linear Attention M., McClelland, J

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.071591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.071591Z digest=sha256:27e46f0b0184329c04ec00ea4cb3a5ba56b68a9913e635165301f17ed385fcaf

Observation ff51e976-4af7-43cb-aa6e-b0a7a000bf31 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

Training Dynamics of In-Context Learning in Linear Attention Linear transformers are secretly fast weight programmers

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.172011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.075725Z digest=sha256:eeed2f74987d2ba8339c5d1b7d2fdeeaaf34afdfb3561fbd6619bbe2ff6969ff

Observation 180988dd-578a-4831-8349-f2b034b461ba · outbound

This paper cites Exponential convergence time of gradient descent for one-dimensional deep linear neural networks.

Training Dynamics of In-Context Learning in Linear Attention Exponential convergence time of gradient descent for one-dimensional deep linear neural networks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.157648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.079454Z digest=sha256:fba7b508a2e8da9012464e4bb9a2f81a91a34689af74b5dd8684c4c198ddfaee

Observation ee85c82c-8ac2-43be-bc0f-a5588a800f5c · outbound

This paper cites Implicit Regularization of Gradient Flow on One-Layer Softmax Attention.

Training Dynamics of In-Context Learning in Linear Attention Implicit Regularization of Gradient Flow on One-Layer Softmax Attention

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.083522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.083522Z digest=sha256:27b42cc9ab256ef0aa94dce4f66b373ab4e4a8ddd48cb3a7b5468e2b8e13d3c1

Observation c52dee80-a090-4389-a912-04517e84c05c · outbound

This paper cites K., Chan, S., Moskovitz, T., Grant, E., Saxe, A., and Hill, F.

Training Dynamics of In-Context Learning in Linear Attention K., Chan, S., Moskovitz, T., Grant, E., Saxe, A., and Hill, F

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.142031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.087876Z digest=sha256:6326b5a416d088a82678d5a8561dd7a9028896d00a7cee7f53c3e0b85da8cd58

Observation 523c51a5-308b-4e46-8576-562225e5584e · outbound

This paper cites K., Moskovitz, T., Hill, F., Chan, S.

Training Dynamics of In-Context Learning in Linear Attention K., Moskovitz, T., Hill, F., Chan, S

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.126088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.091488Z digest=sha256:9822b6b52a838411752e9a01788a4774d56207a9759a4681e2e9f96bc0943c0a

Observation 07cd9722-8346-469b-967c-fcb945bbb1b8 · outbound

This paper cites Strategy Coopetition Explains the Emergence and Transience of In-Context Learning.

Training Dynamics of In-Context Learning in Linear Attention Strategy Coopetition Explains the Emergence and Transience of In-Context Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.095341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.095341Z digest=sha256:f8694cd5257cb2bbd81a5d79dc77071d7f24c3255970f6e765f20a4e094932d2

Observation 5ab6d677-67c2-4c8d-9891-666ecc4bec0d · outbound

This paper cites Unraveling the gradient descent dynamics of transformers.

Training Dynamics of In-Context Learning in Linear Attention Unraveling the gradient descent dynamics of transformers

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.111542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.099630Z digest=sha256:5bf368c139b96654e753e17b2e6c0936bd4f23f1af260983227fbb3a1af43ce4

Observation 17e09882-94d7-458b-9876-e8d11e230fe2 · outbound

This paper cites Transformers as Support Vector Machines.

Training Dynamics of In-Context Learning in Linear Attention Transformers as Support Vector Machines

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.103451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.103451Z digest=sha256:551c546c5cc2b059d894909bf8c060e3a30f6b4e42d165940158903170501c2b

Observation 56750efc-d9f1-4715-8348-b2a22df7b4c7 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.096926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.108132Z digest=sha256:e1395f206f6b21b0839bad7ce4469c0980bdace879a597f7d897bda30764d993

Observation 42d6f779-6a0d-4d90-bc33-edfd1c8735b6 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.082014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.112235Z digest=sha256:e469dd5e74f0b80a45c89a53f0a91649a687ff7e91e5d0da48551e06c74fdec0

Observation a014db2b-f462-4e30-b475-95cf5671a9f3 · outbound

This paper cites Implicit bias and fast convergence rates for self-attention.

Training Dynamics of In-Context Learning in Linear Attention Implicit bias and fast convergence rates for self-attention

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.067318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.116517Z digest=sha256:bb6250a993fda3f0ca6cbc4d3c8ded2946257e862c49b96dcf8b3bedda7c9e80

Observation 4e9d17f5-57a0-432e-918c-b2f0b4f25ebc · outbound

This paper cites N., Kaiser, L.

Training Dynamics of In-Context Learning in Linear Attention N., Kaiser, L

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.120780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.120780Z digest=sha256:4f2ef6903599bffbe9ad3d7d97a6a44c12793336325b7582b99b31779ba87082

Observation 795111a3-fe31-4552-91bb-7dee48da91a3 · outbound

This paper cites Linear transformers are versatile in-context learners.

Training Dynamics of In-Context Learning in Linear Attention Linear transformers are versatile in-context learners

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.125061Z digest=sha256:a4ddf510d2a63e82548e0419bc195e20da0e25cd05c4c9e921a64be97de47b56

Observation 4b478dff-4ebe-4e02-a8bd-9fb779a69a8d · outbound

This paper cites Transformers learn in-context by gradient descent.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn in-context by gradient descent

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.027837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.129291Z digest=sha256:76e197043f3647a54dd877d8cbe583a13908bef63be1541de843695659d2d52f

Observation 26249c99-a298-4db3-be81-cc6fb5aae8bd · outbound

This paper cites How Transformers Get Rich: Approximation and Dynamics Analysis.

Training Dynamics of In-Context Learning in Linear Attention How Transformers Get Rich: Approximation and Dynamics Analysis

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.133732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.133732Z digest=sha256:59c4b77d2b623b8f6ec9bbaae882cc0e1086890ca61315e5bb5b2bdf27e376ca

Observation 24b59c78-541d-4819-8870-0e40b0d80fa7 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.011623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.138237Z digest=sha256:4a758896ce212f2d5f87a9eeb3d1ecc8ea43aee570fb29ff8f9d0de0cc0aaf34

Observation 57f62c0c-2855-4b81-885b-8ab043f8be30 · outbound

This paper cites D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N.

Training Dynamics of In-Context Learning in Linear Attention D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.142516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.142516Z digest=sha256:e61c806a382224011a490672ba8bdc9b650b52c77bf8db679133c96cbd95bdd4

Observation 8d372d5e-e8f2-4ecc-ab0c-cc721c8e3105 · outbound

This paper cites How many pretraining tasks are needed for in-context learning of linear regression? In The Twelfth International Conference on Learning Representations, 2024.

Training Dynamics of In-Context Learning in Linear Attention How many pretraining tasks are needed for in-context learning of linear regression? In The Twelfth International Conference on Learning Representations, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:09.987294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.146831Z digest=sha256:92d62e5af49d8a0fcdba41321bae2522073550cafced9ba626c9f45a7c5b2997

Observation d0ba08e5-4b83-4754-9ab3-53727ac93825 · outbound

This paper cites V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V.

Training Dynamics of In-Context Learning in Linear Attention V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.151276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.151276Z digest=sha256:e4152c855e5988ef48ce903d64581fa2603b5898ab9602f209f09d4d2d2fc9a4

Observation d444a150-a108-4501-a902-0acad08361ed · outbound

This paper cites B., Jegelka, S., and Andreas, J.

Training Dynamics of In-Context Learning in Linear Attention B., Jegelka, S., and Andreas, J

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.155908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.155908Z digest=sha256:7b1b6fce7cc1a50cde34806cd3e3ede2533e8960bfece78d58a2d6c5787e41c7

Observation 1aa313d2-72df-4055-8cd2-343c9770ef16 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:09.972758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.160103Z digest=sha256:10ea90b39d18a12e4eecb4c1f5bb2d5527f0e16869cd4ab3495075d0e6029897

Observation 0cd92c93-764c-431c-912f-57814ddfc26e · outbound

This paper cites In-context learning of a linear transformer block: Benefits of the mlp component and one-step gd initialization.

Training Dynamics of In-Context Learning in Linear Attention In-context learning of a linear transformer block: Benefits of the mlp component and one-step gd initialization

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:09.959015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.164276Z digest=sha256:794030e097dbcc92523a0273a96bfe6d88d5dddf6cad30f7353a2e24fa3ee776

Pith citing papers

Observation 6e666921-58ab-4998-a320-fd8bc5ca6c26 · inbound

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge cites this paper.

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge Training Dynamics of In-Context Learning in Linear Attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:19.125912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:19.125912Z digest=sha256:e27663dffcb9436f967c1390ba5a8929b292b484f8173e5cfed044b3b8079ceb

Observation b2b3f7b8-b40d-4a3c-a664-bb62f0a15515 · inbound

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge cites this paper.

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Training Dynamics of In-Context Learning in Linear Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:44.603104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:44.603104Z digest=sha256:417024e0370b9dd617196a3677c8f0efdc8d9e55f6233bfd1f6eca21bf560726

Observation 7b148a72-2629-48a4-97be-4e1b7a09b141 · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis Training Dynamics of In-Context Learning in Linear Attention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.267015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:54.899039Z digest=sha256:6ed2f4758fef5fd7d4a7c38ffa395e6dac88a08c87b3fda8bd4969007edd1ff7

Observation 66609f14-4854-4d82-97a6-bf1bac3c3e26 · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis Training Dynamics of In-Context Learning in Linear Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:49:44.727751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:49:44.727751Z digest=sha256:188c33ffe6579e16a089abe569d0d1787d1a90579033e61957d481e1cdab1a4e

Observation c0c74abb-4ee8-423a-a6c1-7f2270a1226d · inbound

Specialization of softmax attention heads: insights from the high-dimensional single-location model cites this paper.

Specialization of softmax attention heads: insights from the high-dimensional single-location model Training Dynamics of In-Context Learning in Linear Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:04:54.938955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:04:54.938955Z digest=sha256:d6729fa9c825ed68035be374d3bc9a557f74a03b98401adde6052b1ee81732be

Observation c900c87b-9bb0-439f-b446-83da4b72e296 · inbound

Learning to Adapt: In-Context Learning Beyond Stationarity cites this paper.

Learning to Adapt: In-Context Learning Beyond Stationarity Training Dynamics of In-Context Learning in Linear Attention

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:59.349405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T16:30:36.771589Z digest=sha256:7841ba1664c366e426e0346e5e4fc1ed01775d4b53524d549d02e232aeaab47e

Observation c374f714-ce4e-4396-a6c9-4ed2fc64658a · inbound

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression cites this paper.

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression Training Dynamics of In-Context Learning in Linear Attention

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.409505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T03:12:49.334507Z digest=sha256:dd95a117f51fd97c5e5a3b5c2a26aed268056a3da9d3790ce53c0166e60f51ff

Observation 5de7d0dd-6fd1-4589-a3c3-365faa864a81 · inbound

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression cites this paper.

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression Training Dynamics of In-Context Learning in Linear Attention

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:09.452218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T22:26:55.610155Z digest=sha256:80a5d5151aca19cdd76333daf9b4fa4b4ee25865909eed416f1bba569f61127d

Observation d81efff2-39de-4965-b583-05dda0263c3f · inbound

An Asymptotic Theory of Chain-of-Thought in In-Context Learning cites this paper.

An Asymptotic Theory of Chain-of-Thought in In-Context Learning Training Dynamics of In-Context Learning in Linear Attention

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:38.978513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T08:25:37.421049Z digest=sha256:bd4fc1eb9dfc24f31e05e0657d392f4096c7d4d7c4e9a492b5fd5b5c02e31697

Observation 878a2b8a-6003-46f5-97a5-066f7bbd83b7 · inbound

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway cites this paper.

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway Training Dynamics of In-Context Learning in Linear Attention

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:46.412075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T23:15:59.394256Z digest=sha256:8789ae9880551df8f3aa6cf277e45e25b4f93617c2b4246eaf5a70afeb939ca1

Observation de61a2c5-ec38-437d-ad1b-38ab5e36bc22 · inbound

Towards personalised intervention: A causal-dynamical framework to determine psychological treatment trajectories cites this paper.

Towards personalised intervention: A causal-dynamical framework to determine psychological treatment trajectories Training Dynamics of In-Context Learning in Linear Attention

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:36.731429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T14:13:28.981530Z digest=sha256:39e9e2a4d8ce4463baa3c6131e18a964ada88f50e69467543e611b0e87b403eb

Observation 044ed462-b13a-40e1-85df-726ac21d695b · inbound

Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch cites this paper.

Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch Training Dynamics of In-Context Learning in Linear Attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T00:51:41.905788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:51:41.905788Z digest=sha256:7a80197a0da75486e80afb3c197f72c3fc08b33dfcf84371675ba2fedbbdce47