Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics of In-Context Learning in Linear Attention

As of 17 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 12 inbound Pith citation observations for arXiv:2501.16265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16265 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:42:09.164276Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:54:19.125912Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 863a75de-3fea-426f-beb9-9949540cbda1 · outbound

This paper cites write newline.

Training Dynamics of In-Context Learning in Linear Attention write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.771564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.771564Z digest=sha256:98af1b178726a9cb8ec6e79dd89a9cee393849218cffb548b8c15b99678014d5

Observation 72351a8a-da9b-4801-a986-b92766eb20b9 · outbound

This paper cites Context-Scaling versus Task-Scaling in In-Context Learning.

Training Dynamics of In-Context Learning in Linear Attention Context-Scaling versus Task-Scaling in In-Context Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-10T13:42:09.943815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.777812Z digest=sha256:3726660189818b805cf129d1f3b2d79fc4a36f61fce9be56aa53c54940b4d464

Observation 64959c77-d4e0-4ead-8c55-ed7fb68768ca · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.784728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.784728Z digest=sha256:ece561062c93c886a4f7b5a0aa2f370e372ba4f02992e41cd3e681313465a187

Observation 0e7d4324-81d1-4fc7-8515-0dcccf45abe0 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Training Dynamics of In-Context Learning in Linear Attention Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.789849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.789849Z digest=sha256:809ae2e38699fe1b88a3330c41af0c2d27e9d0c7b0c68fe1e33fc65096ef477d

Observation e7b2f64c-7709-4c3f-b47e-b65656303604 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

Training Dynamics of In-Context Learning in Linear Attention What learning algorithm is in-context learning? investigations with linear models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.794517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.794517Z digest=sha256:eb4576c51f6f4a7309b05075848a1a2c8a363482640cb80a0845eb930eecdf2d

Observation d12b03fd-e8ef-410e-9551-b1ffaaaf4593 · outbound

This paper cites A., Merullo, J., and Pavlick, E.

Training Dynamics of In-Context Learning in Linear Attention A., Merullo, J., and Pavlick, E

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.799444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.799444Z digest=sha256:c7e8109256e4262598fdda89027f090d70108e0813155296a029c1903f06ec82

Observation 5d7f1b81-2cdc-4b17-a1af-749a4644f8a3 · outbound

This paper cites Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens.

Training Dynamics of In-Context Learning in Linear Attention Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.805009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.805009Z digest=sha256:e3f6d3b731a01cc45482daab1ff5415fea4a2c6fc7c50869e999db003e30fe74

Observation f7074a8a-00f3-4479-9263-5718ab8df39f · outbound

This paper cites A convergence analysis of gradient descent for deep linear neural networks.

Training Dynamics of In-Context Learning in Linear Attention A convergence analysis of gradient descent for deep linear neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.852752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.811298Z digest=sha256:51b03786e7cb33d309a567f71387da5868718458e56fdd095338fce5dad900b0

Observation 5c3b3432-1ef3-4765-aacf-9aef3aad3703 · outbound

This paper cites Max-margin token selection in attention mechanism.

Training Dynamics of In-Context Learning in Linear Attention Max-margin token selection in attention mechanism

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.838845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.815984Z digest=sha256:539c28cd142f712598ca05dcf7391edb3d0c786a20812df6eae62fc0b5abcd2c

Observation 9cf46489-6ab3-41b9-bc27-ae77d20a4f8a · outbound

This paper cites Neural networks as kernel learners: The silent alignment effect.

Training Dynamics of In-Context Learning in Linear Attention Neural networks as kernel learners: The silent alignment effect

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.820551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.820551Z digest=sha256:eb16f6534c00f0a4d0603a34095a9c5a07eaa230c05aa0b77461bc637f4ff096

Observation 349988df-dfb1-4152-9151-a32999870656 · outbound

This paper cites Transformers as statisticians: Provable in-context learning with in-context algorithm selection.

Training Dynamics of In-Context Learning in Linear Attention Transformers as statisticians: Provable in-context learning with in-context algorithm selection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.815157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.825679Z digest=sha256:46dbc9acbcf17c1f164938815c39b8773fdf1567c209ce2dfc2e699f8b6852ab

Observation bb2831a1-992f-4bca-9856-b729fbcd9d85 · outbound

This paper cites Transformers learn through gradual rank increase.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn through gradual rank increase

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.800154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.830426Z digest=sha256:3c8dc553c2c462622a4e785fc38979e83004c471f89d078d3556393a572d8a78

Observation 66622513-0ceb-48fe-b15f-ce4584ee3eba · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.784524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.835092Z digest=sha256:55f124c7c59c511d1becfb63d99ec92b8f9b6744639f64a142ce91ffef01bb0d

Observation a49640b7-93d8-485e-a725-2fe526ba29fd · outbound

This paper cites Infinite limits of multi-head transformer dynamics.

Training Dynamics of In-Context Learning in Linear Attention Infinite limits of multi-head transformer dynamics

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.770121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.839610Z digest=sha256:15d35c0a7cc6d405eb4ade23413d283bf78002843990863f05600313cb1eb044

Observation 4b316124-1a7f-4e01-a6df-cee59b3afa68 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.844836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.844836Z digest=sha256:9bf5cc7ee40f9fe00510228b920fd4a01e1a76f51f38257be20d19c1bef1fa27

Observation 9431fafa-7de8-4929-b83c-34a635225f91 · outbound

This paper cites Toward understanding in-context vs.

Training Dynamics of In-Context Learning in Linear Attention Toward understanding in-context vs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.745978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.849357Z digest=sha256:11986ce8b68573d00106350af624b7bab5b5437991bb6617d0412c7162b3d47a

Observation deec83fe-93d7-433b-9baf-a99c3158bce9 · outbound

This paper cites Data distributional properties drive emergent in-context learning in transformers.

Training Dynamics of In-Context Learning in Linear Attention Data distributional properties drive emergent in-context learning in transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.731984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.854192Z digest=sha256:ddc4f20c1c98294e1d23070399a854298e0b0d527ef1dc047d6d2d56c66e10dd

Observation cdd4c5b0-3ed7-421d-bff1-ef95b33764c2 · outbound

This paper cites L., and Saphra, N.

Training Dynamics of In-Context Learning in Linear Attention L., and Saphra, N

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.718066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.858728Z digest=sha256:9bc3dbb0374854d9f065f6b43732e492570871dfc084a96ebe44c89f0f69b302

Observation 51982e89-d015-491b-8279-7a1ff550281b · outbound

This paper cites Provably learning a multi-head attention layer.

Training Dynamics of In-Context Learning in Linear Attention Provably learning a multi-head attention layer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.863656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.863656Z digest=sha256:a67c3a897c7932ae4ba57884039107e27d2edd86545b9de50259988191da92bb

Observation 48c82c41-46f1-484f-9f4d-d176ffbd6535 · outbound

This paper cites Training dynamics of multi-head softmax attention for in-context learning: Emergence, convergence, and optimality.

Training Dynamics of In-Context Learning in Linear Attention Training dynamics of multi-head softmax attention for in-context learning: Emergence, convergence, and optimality

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.704119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.868864Z digest=sha256:0f01d5253a2e1b949775b7e8dce99086d9ef02ca21f71012ffe81e3f5d99497e

Observation 2dfe84c6-4a23-4767-90aa-9e8e1a1b41e7 · outbound

This paper cites Unveiling induction heads: Provable training dynamics and feature learning in transformers.

Training Dynamics of In-Context Learning in Linear Attention Unveiling induction heads: Provable training dynamics and feature learning in transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.689327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.873299Z digest=sha256:9863a55cad94b58d706d1d445a9ed6bdbdd0d40fa066eed898d3723d16abf069

Observation 8e394567-653e-4201-9626-cbb173be8e5a · outbound

This paper cites On lazy training in differentiable programming.

Training Dynamics of In-Context Learning in Linear Attention On lazy training in differentiable programming

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.674486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.877770Z digest=sha256:f93132b87409dfc01a8c88b01e9b5c05d69add1e3f29dcb3fb38e9a9f8fcd332

Observation e00d3fd6-135a-4b81-932c-f81f3fd0b8bb · outbound

This paper cites S., Hu, W., and Lee, J.

Training Dynamics of In-Context Learning in Linear Attention S., Hu, W., and Lee, J

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.659493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.882860Z digest=sha256:48f4932b89107e9f13f4bf67a53b67e21bbce462ed3382a2f9bc765cee759351

Observation eb1251fe-4819-42d0-b142-b4df18a1c9d9 · outbound

This paper cites Finite Sample Analysis and Bounds of Generalization Error of Gradient Descent in In-Context Linear Regression.

Training Dynamics of In-Context Learning in Linear Attention Finite Sample Analysis and Bounds of Generalization Error of Gradient Descent in In-Context Linear Regression

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.887996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.887996Z digest=sha256:af4f9e8c45f8260cd4ba2d1dd03d348a2aed4ad4fee1a3232495b6139e1c2784

Observation e1295f8f-2108-4b23-8d35-5584f6f90474 · outbound

This paper cites The evolution of statistical induction heads: In-context learning markov chains.

Training Dynamics of In-Context Learning in Linear Attention The evolution of statistical induction heads: In-context learning markov chains

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.646077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.893622Z digest=sha256:220dd7315edc4d9ff606e118dbb1a1658b9cb40ffb53f414789d5ce12933ce27

Observation 21e847e9-0a87-4176-813d-e0ee1a1e4fb8 · outbound

This paper cites and Vardi, G.

Training Dynamics of In-Context Learning in Linear Attention and Vardi, G

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.630905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.898470Z digest=sha256:e512b5c888b565600ac0f6fd6934081710f727b194278c88afeb570789dadf5d

Observation 71849b5f-4342-4594-9571-5e9f60e61493 · outbound

This paper cites Transformers learn to achieve second-order convergence rates for in-context linear regression.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn to achieve second-order convergence rates for in-context linear regression

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.617623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.903611Z digest=sha256:2d82f7fed9cf2002bca00887580a404b0e8e4705e31a308ec9c6ba19786cfa99

Observation 5fce92da-7176-4c50-a44b-b8e47b0f7504 · outbound

This paper cites Effect of batch learning in multilayer neural networks.

Training Dynamics of In-Context Learning in Linear Attention Effect of batch learning in multilayer neural networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.602828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.908134Z digest=sha256:049e66850b4d312771f13a852c1963e05a87dda2df5db05839327b11e4829794

Observation d3717c82-e22d-4d4a-a061-78241bdef035 · outbound

This paper cites S., and Valiant, G.

Training Dynamics of In-Context Learning in Linear Attention S., and Valiant, G

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.587295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.912602Z digest=sha256:56a53237daf1ec8f66ecdd2a5f1f1f0e578b84a33f25d0241d6b6b15ac8b4ade

Observation 22b504e4-4ff0-43fe-b746-85863551cc3e · outbound

This paper cites On the Role of Depth and Looping for In-Context Learning with Task Diversity.

Training Dynamics of In-Context Learning in Linear Attention On the Role of Depth and Looping for In-Context Learning with Task Diversity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.917150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.917150Z digest=sha256:a7009b94dc56abb46a7c598372486018dff4d39b9091203611fc38dc2e9871fd

Observation 080117ce-e8b2-4bcb-b8ca-b821c4b05ab7 · outbound

This paper cites Dynamic metastability in the self-attention model.

Training Dynamics of In-Context Learning in Linear Attention Dynamic metastability in the self-attention model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.922647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.922647Z digest=sha256:7356ca1588d025e824cd38facb4a4e6b7b7dbd28e6bb215316b663372ae92f80

Observation a1fc2764-3e10-4b0f-8f81-e906ccb50c9d · outbound

This paper cites In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention.

Training Dynamics of In-Context Learning in Linear Attention In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.927759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.927759Z digest=sha256:6ece3465714664d9227279d37dd4f99a8858c0e66993fddf29618cc35868f902

Observation 6860ba27-79d8-4f9a-b077-ae624ebe29f6 · outbound

This paper cites Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks.

Training Dynamics of In-Context Learning in Linear Attention Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.572786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.934151Z digest=sha256:f54998f18cebf1e26899854df19350acb96217ba7a35eb4df7126c2cf5caab28

Observation 0e993c53-abc7-4b3a-84a4-c351153cc72a · outbound

This paper cites T., Schrodi, S., Bratuli\' c , J., Behrmann, N., Fischer, V., and Brox, T.

Training Dynamics of In-Context Learning in Linear Attention T., Schrodi, S., Bratuli\' c , J., Behrmann, N., Fischer, V., and Brox, T

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.557977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.938582Z digest=sha256:4e3b48a50cdf4e7eca7c7768e4cd45fea058eda52bb234a59ed6a0d7bbb95156

Observation acf3c34c-7f88-4ad2-98a2-954ff45b44d5 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.543221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.943036Z digest=sha256:8716435e72f7e812b8ac393e1474f344883a5a3adfc193bcd0c394da50606d1c

Observation 7ecd2b09-c6dc-49c3-ba91-125b688ed482 · outbound

This paper cites Non-asymptotic convergence of training transformers for next-token prediction.

Training Dynamics of In-Context Learning in Linear Attention Non-asymptotic convergence of training transformers for next-token prediction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.528891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.947058Z digest=sha256:6e2db7625903c170213c712edd425d847259e676fdd8fbd2dbc82a4992793dcd

Observation f74697cf-cb95-4040-afec-0b0370f1f431 · outbound

This paper cites In-context convergence of transformers.

Training Dynamics of In-Context Learning in Linear Attention In-context convergence of transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.513722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.951456Z digest=sha256:b5328645a05aac0ae498e7da05f961a90feba740038c612b5ddf2b86dac57fb3

Observation 7d00bd3f-797c-4623-93a3-701fea851152 · outbound

This paper cites A theoretical analysis of self-supervised learning for vision transformers.

Training Dynamics of In-Context Learning in Linear Attention A theoretical analysis of self-supervised learning for vision transformers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.499288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.956251Z digest=sha256:6e70c063c34a03337d8258f00ed6b76e97e07af5c607c7ffa5e17e135c13ded5

Observation a1af68c1-b9e9-4290-a60a-198e97f63f17 · outbound

This paper cites E., Huang, Y., Li, Y., Rawat, A.

Training Dynamics of In-Context Learning in Linear Attention E., Huang, Y., Li, Y., Rawat, A

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.483227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.960781Z digest=sha256:4f88dde1388e5c1a71aa86b11cdede5f99c730f1fd4307b739807e9ebd9e1f2c

Observation 94160fa6-dc10-40e2-96d4-fee0712934d6 · outbound

This paper cites D., and Ryu, E.

Training Dynamics of In-Context Learning in Linear Attention D., and Ryu, E

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.462332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.965195Z digest=sha256:b54ebae2cbb1203d3ce15d56b008b3f4cf6fa6276a931b11f8c3d03a80ab3951

Observation d8050a11-f76c-47c8-a94a-6973553b1eec · outbound

This paper cites Vision transformers provably learn spatial structure.

Training Dynamics of In-Context Learning in Linear Attention Vision transformers provably learn spatial structure

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.446003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.969754Z digest=sha256:7c3fe9995b11a068db320f77ae589f363bd73a52c7c5a5051014e88a3afe9e74

Observation d3c3ecfa-dfca-41db-89e8-8b638a718ee0 · outbound

This paper cites and Telgarsky, M.

Training Dynamics of In-Context Learning in Linear Attention and Telgarsky, M

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.430027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.974419Z digest=sha256:a8cfda685b0ed20671da7e4cf94273e6f786577ea0ec1e153f963f07cd3e5289

Observation a5dff65c-27b7-4a06-8c01-740afd2959a5 · outbound

This paper cites Unveil benign overfitting for transformer in vision: Training dynamics, convergence, and generalization.

Training Dynamics of In-Context Learning in Linear Attention Unveil benign overfitting for transformer in vision: Training dynamics, convergence, and generalization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.412050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.978996Z digest=sha256:b9afe059ebab6224ce26757c6de43df5fbba31d5b03b3e6a99e740425f7528dc

Observation 91ed1275-8bee-458e-8065-2b87698e280a · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.983616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.983616Z digest=sha256:00f5c2c332e112542e6d9da3e9980f489f48df861d29415b3022bdbb2cb5bab3

Observation 050d0fdb-a8db-43ce-a50f-a50a63825831 · outbound

This paper cites and Suzuki, T.

Training Dynamics of In-Context Learning in Linear Attention and Suzuki, T

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.394142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:08.988243Z digest=sha256:c63b866e71454f2fe63d80a98d6664801b32c370a7655f7d0882ef879e916d20

Observation 5b6df4c7-96ab-467b-b6ef-96bda7b3e719 · outbound

This paper cites Geometry of linear convolutional networks.

Training Dynamics of In-Context Learning in Linear Attention Geometry of linear convolutional networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.992917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.992917Z digest=sha256:3891780d1bc6db6a903780e039d72b1924b76eab34947bb40e25b37c30057b1d

Observation 7adbada2-0e9e-4892-9d7d-97f8fc51d6cd · outbound

This paper cites Function space and critical points of linear convolutional networks.

Training Dynamics of In-Context Learning in Linear Attention Function space and critical points of linear convolutional networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.997544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.997544Z digest=sha256:e02ce4d9d6375ac8b69843952431609ee0acb08766140a2d806c643fdeb7597e

Observation ac224662-0663-4de0-ad94-82ec03472059 · outbound

This paper cites Is attention required for ICL ? exploring the relationship between model architecture and in-context learning ability.

Training Dynamics of In-Context Learning in Linear Attention Is attention required for ICL ? exploring the relationship between model architecture and in-context learning ability

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.377801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.002949Z digest=sha256:81164018587a4f6a30cd4ff2e68e0b3f5d388c3235b3e99570127869bf9bfdce

Observation 1c9dc810-147c-4185-8fc3-e45f9cee7c70 · outbound

This paper cites Fine-grained analysis of in-context linear estimation: Data, architecture, and beyond.

Training Dynamics of In-Context Learning in Linear Attention Fine-grained analysis of in-context linear estimation: Data, architecture, and beyond

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.361926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.007960Z digest=sha256:5a50d104aa1a922874adc40679e91ab32c84247533087ea83e7b656fb14c3cfb

Observation bfcc9474-cd07-44c3-a664-9f1530f7c447 · outbound

This paper cites M., Letey, M.

Training Dynamics of In-Context Learning in Linear Attention M., Letey, M

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.012085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.012085Z digest=sha256:1c96e0917fb5f01a93f85a1ab41bc677015bfc2884c45c8e1149cd20ec6a87e5

Observation fb17ad90-e3da-42b1-9f02-1f4e87fa2969 · outbound

This paper cites V., Hashimoto, T., and Ma, T.

Training Dynamics of In-Context Learning in Linear Attention V., Hashimoto, T., and Ma, T

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.345757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.016999Z digest=sha256:2754c90b8a5d17e4b57fa5a50355657f4c20c94adf468c0978aaf8bfe76e4da4

Observation 35bb2cd7-dc34-417c-ad0d-6a6a4981f5cc · outbound

This paper cites V., Bondaschi, M., Girish, A., Nagle, A., Kim, H., Gastpar, M., and Ekbote, C.

Training Dynamics of In-Context Learning in Linear Attention V., Bondaschi, M., Girish, A., Nagle, A., Kim, H., Gastpar, M., and Ekbote, C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.328119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.022056Z digest=sha256:cd170101ad7cfe9d1bee2e7218a24b10015fa1185232b9273fa6ec5b3173a08c

Observation 5b4a75ca-26e7-4ecb-9977-4d41283da5f6 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Training Dynamics of In-Context Learning in Linear Attention Progress measures for grokking via mechanistic interpretability

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.026844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.026844Z digest=sha256:824a6e66d4c3fd38073b764b1822cb0944dd16fdfb282fea35afb121e2b86ab0

Observation daa75e13-4970-47e2-8c0e-02553d455f70 · outbound

This paper cites and Reddy, G.

Training Dynamics of In-Context Learning in Linear Attention and Reddy, G

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.299913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.031612Z digest=sha256:bbed29bb74e288e887e2519b0a4ae47336bb9b3f8b10c03dae3551741408019f

Observation 927c95b6-68bb-4d4e-91d3-68fda00b5b24 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.280694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.036312Z digest=sha256:5a33fb8d6e134d83b6337afe367d5ba3f4a339da82435d9ae651c1d8fd718e03

Observation 98433b6e-25fe-4c04-a3e6-b8278de044cb · outbound

This paper cites In-context learning and induction heads, 2022.

Training Dynamics of In-Context Learning in Linear Attention In-context learning and induction heads, 2022

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.263268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.040604Z digest=sha256:8724a40b9b33c3ac4e032fee3fc9e61925c7a61fdcd53f6280c1d2a03a220b4c

Observation 161f5a43-37d3-47d4-bb9c-813d4208ea58 · outbound

This paper cites and Reznikoff, M.

Training Dynamics of In-Context Learning in Linear Attention and Reznikoff, M

Reference 57

Resolution
verified exact
doi, observed 2026-08-10T13:42:09.225530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.044925Z digest=sha256:192ccc9132fcf74e47ca2e11e458876180afe0b43358de486a46d15a6d59c163

Observation ef1435cf-fe21-443d-a782-3920db0624bc · outbound

This paper cites F., Lubana, E.

Training Dynamics of In-Context Learning in Linear Attention F., Lubana, E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.240533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.049180Z digest=sha256:b92096949f62fde7f39d314373993896b5184551d3f448d01c1ec76836e2be0e

Observation e69add26-a54c-4ff5-98e6-f969c065b3df · outbound

This paper cites The mechanistic basis of data dependence and abrupt learning in an in-context classification task.

Training Dynamics of In-Context Learning in Linear Attention The mechanistic basis of data dependence and abrupt learning in an in-context classification task

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.053680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.053680Z digest=sha256:f8a2a5a7d61f6cd2ade12e4da341ee2adc76103f422664f92631af553c0218be

Observation f49ea934-b4d3-490d-a612-247e02399420 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.213412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.058068Z digest=sha256:7e022017309036a8b417ef7a4ac11c3b8b08e6742eeb8a1aafcceba6ad91696f

Observation 2d75092c-e492-4936-8c93-c2574e5815d2 · outbound

This paper cites A distributional simplicity bias in the learning dynamics of transformers.

Training Dynamics of In-Context Learning in Linear Attention A distributional simplicity bias in the learning dynamics of transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.199875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.063049Z digest=sha256:6b31e3c9ae98c81d80336c7d6d8c3e042889244ac484986e4c5601aba1eb5404

Observation 33e03add-e130-498c-aabe-26e79c5e8aa7 · outbound

This paper cites M., McClelland, J.

Training Dynamics of In-Context Learning in Linear Attention M., McClelland, J

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.186428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.067331Z digest=sha256:90e0d09187011c11d36afadfdabefa93a0b6cb9fd6d558b0aff62239094c8b3c

Observation 0ae0ce06-93b8-44f8-85f5-afdeba87eb8f · outbound

This paper cites M., McClelland, J.

Training Dynamics of In-Context Learning in Linear Attention M., McClelland, J

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.071591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.071591Z digest=sha256:fcd7e2489ea92a7c02153a01b2af3162b9c7e6060bfb5caec7ac11e629288dbd

Observation ff51e976-4af7-43cb-aa6e-b0a7a000bf31 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

Training Dynamics of In-Context Learning in Linear Attention Linear transformers are secretly fast weight programmers

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.172011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.075725Z digest=sha256:9fdde59759f13084296fdc8f3fbfd85f9f14c73c4273123ca16c387b5bd27d94

Observation 180988dd-578a-4831-8349-f2b034b461ba · outbound

This paper cites Exponential convergence time of gradient descent for one-dimensional deep linear neural networks.

Training Dynamics of In-Context Learning in Linear Attention Exponential convergence time of gradient descent for one-dimensional deep linear neural networks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.157648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.079454Z digest=sha256:1ddae67f496d8268517b4e81d24cf270815a076a98614e914d8735118c11ca2e

Observation ee85c82c-8ac2-43be-bc0f-a5588a800f5c · outbound

This paper cites Implicit Regularization of Gradient Flow on One-Layer Softmax Attention.

Training Dynamics of In-Context Learning in Linear Attention Implicit Regularization of Gradient Flow on One-Layer Softmax Attention

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.083522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.083522Z digest=sha256:4a0835fa5033a072f4b5a5db2eafa07865c6068c33567cdcad09f84309955849

Observation c52dee80-a090-4389-a912-04517e84c05c · outbound

This paper cites K., Chan, S., Moskovitz, T., Grant, E., Saxe, A., and Hill, F.

Training Dynamics of In-Context Learning in Linear Attention K., Chan, S., Moskovitz, T., Grant, E., Saxe, A., and Hill, F

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.142031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.087876Z digest=sha256:c28226393fc69a8c439ac98de528a49d69866f85b60317f24283767860ec37ba

Observation 523c51a5-308b-4e46-8576-562225e5584e · outbound

This paper cites K., Moskovitz, T., Hill, F., Chan, S.

Training Dynamics of In-Context Learning in Linear Attention K., Moskovitz, T., Hill, F., Chan, S

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.126088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.091488Z digest=sha256:152b24aa16ab2cb23d190e3542550196bee4df2a0cdcf05a5b50ad599163b61f

Observation 07cd9722-8346-469b-967c-fcb945bbb1b8 · outbound

This paper cites Strategy Coopetition Explains the Emergence and Transience of In-Context Learning.

Training Dynamics of In-Context Learning in Linear Attention Strategy Coopetition Explains the Emergence and Transience of In-Context Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.095341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.095341Z digest=sha256:b34a47e082d00dbe216ad2fc59a9a26beb0c9edddcfd2126355e247e56d18292

Observation 5ab6d677-67c2-4c8d-9891-666ecc4bec0d · outbound

This paper cites Unraveling the gradient descent dynamics of transformers.

Training Dynamics of In-Context Learning in Linear Attention Unraveling the gradient descent dynamics of transformers

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.111542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.099630Z digest=sha256:62cc982646648abeee9f65e70a49ddf712186fc696508017ac5caa5673e23276

Observation 17e09882-94d7-458b-9876-e8d11e230fe2 · outbound

This paper cites Transformers as Support Vector Machines.

Training Dynamics of In-Context Learning in Linear Attention Transformers as Support Vector Machines

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.103451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.103451Z digest=sha256:0deb8e7d1e105d610051564d14550f0043b8665963afa6e18c34bf94e339bb45

Observation 56750efc-d9f1-4715-8348-b2a22df7b4c7 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.096926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.108132Z digest=sha256:26aba1a4dadb274b98bc504033eac8eca102adb101b6113b20f2c0c885a790ad

Observation 42d6f779-6a0d-4d90-bc33-edfd1c8735b6 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.082014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.112235Z digest=sha256:f6a8380d38becf52aa94af12c09c6df143b730fba493908bec23ff5ffc8f5325

Observation a014db2b-f462-4e30-b475-95cf5671a9f3 · outbound

This paper cites Implicit bias and fast convergence rates for self-attention.

Training Dynamics of In-Context Learning in Linear Attention Implicit bias and fast convergence rates for self-attention

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.067318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.116517Z digest=sha256:6493438a2685ad8c102510039c43fcaebc44306587561909754a3ae6e08199f6

Observation 4e9d17f5-57a0-432e-918c-b2f0b4f25ebc · outbound

This paper cites N., Kaiser, L.

Training Dynamics of In-Context Learning in Linear Attention N., Kaiser, L

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.120780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.120780Z digest=sha256:30e9a92c9f7f40a518dbf2aae5b26662ed05863fdf15621a046dd281078f8653

Observation 795111a3-fe31-4552-91bb-7dee48da91a3 · outbound

This paper cites Linear transformers are versatile in-context learners.

Training Dynamics of In-Context Learning in Linear Attention Linear transformers are versatile in-context learners

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.125061Z digest=sha256:4afc272dfad7d9a9667694ce7ca03e33de4dfbc547df402c32bf9852c839dc35

Observation 4b478dff-4ebe-4e02-a8bd-9fb779a69a8d · outbound

This paper cites Transformers learn in-context by gradient descent.

Training Dynamics of In-Context Learning in Linear Attention Transformers learn in-context by gradient descent

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:10.027837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.129291Z digest=sha256:2cdf821bdc9436fcbd53df16700aa707ab9f02fb2f71a46b0abd2ffd7e3a8f6d

Observation 26249c99-a298-4db3-be81-cc6fb5aae8bd · outbound

This paper cites How Transformers Get Rich: Approximation and Dynamics Analysis.

Training Dynamics of In-Context Learning in Linear Attention How Transformers Get Rich: Approximation and Dynamics Analysis

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.133732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.133732Z digest=sha256:685d7ceb6b209ca3a2151514d16dcce57b0e0ef7a32359ad259954fae282ffc5

Observation 24b59c78-541d-4819-8870-0e40b0d80fa7 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:10.011623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.138237Z digest=sha256:e93420ac45355f35608ec6fbfd2931871c693b2be70f340aa1c43582c1826c04

Observation 57f62c0c-2855-4b81-885b-8ab043f8be30 · outbound

This paper cites D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N.

Training Dynamics of In-Context Learning in Linear Attention D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.142516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.142516Z digest=sha256:7d7b67c978e2b7be67c65d7b5319cb751c857442d13d5dc507551f38a69dad47

Observation 8d372d5e-e8f2-4ecc-ab0c-cc721c8e3105 · outbound

This paper cites How many pretraining tasks are needed for in-context learning of linear regression? In The Twelfth International Conference on Learning Representations, 2024.

Training Dynamics of In-Context Learning in Linear Attention How many pretraining tasks are needed for in-context learning of linear regression? In The Twelfth International Conference on Learning Representations, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:09.987294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.146831Z digest=sha256:598f4cfef9c00d418c0dc0b9f63c13451674f9e2983fe85812fdc68c5062c75b

Observation d0ba08e5-4b83-4754-9ab3-53727ac93825 · outbound

This paper cites V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V.

Training Dynamics of In-Context Learning in Linear Attention V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.151276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.151276Z digest=sha256:bbf6ed38eacaba6c9d943ab1ed4ff30fd3b9c5ba3a5eacb673ffca15ac5e3a6d

Observation d444a150-a108-4501-a902-0acad08361ed · outbound

This paper cites B., Jegelka, S., and Andreas, J.

Training Dynamics of In-Context Learning in Linear Attention B., Jegelka, S., and Andreas, J

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:09.155908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:09.155908Z digest=sha256:e5bdfff9aca27bae879c66728c10f16a4d0810d4e9b330fb097f825570750c92

Observation 1aa313d2-72df-4055-8cd2-343c9770ef16 · outbound

This paper cites an unresolved cited work.

Training Dynamics of In-Context Learning in Linear Attention Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:42:09.972758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.160103Z digest=sha256:29f70ccdb46ca106c18e4e593353c6c00b36619c8cf89dd1d98a65b5a8cfc8ee

Observation 0cd92c93-764c-431c-912f-57814ddfc26e · outbound

This paper cites In-context learning of a linear transformer block: Benefits of the mlp component and one-step gd initialization.

Training Dynamics of In-Context Learning in Linear Attention In-context learning of a linear transformer block: Benefits of the mlp component and one-step gd initialization

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:42:09.959015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T13:42:09.164276Z digest=sha256:815d55392b7892b198ee4e8b8991db420ae3a48e65d280eaf12cf7f4fe5baeb4

Pith citing papers

Observation 6e666921-58ab-4998-a320-fd8bc5ca6c26 · inbound

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge cites this paper.

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge Training Dynamics of In-Context Learning in Linear Attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:19.125912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:19.125912Z digest=sha256:7ee1baa360aeae57a2d0aa00bf5aff7ef09127865cfa66a277abd8ddbb213e81

Observation b2b3f7b8-b40d-4a3c-a664-bb62f0a15515 · inbound

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge cites this paper.

Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Training Dynamics of In-Context Learning in Linear Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:44.603104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:44.603104Z digest=sha256:61852e3eb3f1d95ce89c6a4216fcd5dfb79228158059d7cca9ab39e06d9f87bb

Observation 7b148a72-2629-48a4-97be-4e1b7a09b141 · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis Training Dynamics of In-Context Learning in Linear Attention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.267015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T18:09:54.899039Z digest=sha256:dc0ea92b785988ee9bbdd50e3abcc1c91bc21bba6ba2ed14f843e5e3c6b1784a

Observation 66609f14-4854-4d82-97a6-bf1bac3c3e26 · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis Training Dynamics of In-Context Learning in Linear Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:49:44.727751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:49:44.727751Z digest=sha256:3604d3bb6ef9d9b3e1e1e9829c3fe2c493646639cddf31012e15ba7a9f4a46bb

Observation c0c74abb-4ee8-423a-a6c1-7f2270a1226d · inbound

Specialization of softmax attention heads: insights from the high-dimensional single-location model cites this paper.

Specialization of softmax attention heads: insights from the high-dimensional single-location model Training Dynamics of In-Context Learning in Linear Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:04:54.938955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:04:54.938955Z digest=sha256:3d9320984c3120a7afc4b30eef96dc193c8345716554bdd4dc452ae02a1b7bc7

Observation c900c87b-9bb0-439f-b446-83da4b72e296 · inbound

Learning to Adapt: In-Context Learning Beyond Stationarity cites this paper.

Learning to Adapt: In-Context Learning Beyond Stationarity Training Dynamics of In-Context Learning in Linear Attention

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:59.349405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T16:30:36.771589Z digest=sha256:e3ea3918f0c33139b02e2a88b82c18b1ac60ae935c59eee1692bb7b4eebe4ecc

Observation c374f714-ce4e-4396-a6c9-4ed2fc64658a · inbound

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression cites this paper.

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression Training Dynamics of In-Context Learning in Linear Attention

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.409505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:12:49.334507Z digest=sha256:2f75331d529bff34a36b5d7e1cb464e53479664841158e08698198c98baf4544

Observation 5de7d0dd-6fd1-4589-a3c3-365faa864a81 · inbound

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression cites this paper.

Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression Training Dynamics of In-Context Learning in Linear Attention

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:09.452218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T22:26:55.610155Z digest=sha256:3c2e77fa1d0b9622e768b8dd82c5f114fdbdd0aeeb308c1471fd0923015e09f9

Observation d81efff2-39de-4965-b583-05dda0263c3f · inbound

An Asymptotic Theory of Chain-of-Thought in In-Context Learning cites this paper.

An Asymptotic Theory of Chain-of-Thought in In-Context Learning Training Dynamics of In-Context Learning in Linear Attention

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:38.978513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T08:25:37.421049Z digest=sha256:d44add9cc27ba028c442a204e8dbba51b82f1aca1a0130b7dcf3860ce5dcc2df

Observation 878a2b8a-6003-46f5-97a5-066f7bbd83b7 · inbound

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway cites this paper.

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway Training Dynamics of In-Context Learning in Linear Attention

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:46.412075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T23:15:59.394256Z digest=sha256:c38b58cb2d4d683c3f902f9359c82f6977d06108a15a375db0ac8a1425be52df

Observation de61a2c5-ec38-437d-ad1b-38ab5e36bc22 · inbound

Towards personalised intervention: A causal-dynamical framework to determine psychological treatment trajectories cites this paper.

Towards personalised intervention: A causal-dynamical framework to determine psychological treatment trajectories Training Dynamics of In-Context Learning in Linear Attention

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:36.731429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T14:13:28.981530Z digest=sha256:709698db10cf2c6a2f352d45c490592051bb8bdfb0f2951df884450bbb36b906

Observation 044ed462-b13a-40e1-85df-726ac21d695b · inbound

Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch cites this paper.

Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch Training Dynamics of In-Context Learning in Linear Attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T00:51:41.905788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:51:41.905788Z digest=sha256:7eb178189737c39af9cd4d9472ffc545614d29d6a4dcdf692837d2f42fbe820b