Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2402.19442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19442 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:39:54.990906Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:42:49.980406Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a7ddefc3-7877-40e9-98f9-c5866659a0df · inbound

Is In-Context Universality Enough? MLPs are Also Universal In-Context cites this paper.

Is In-Context Universality Enough? MLPs are Also Universal In-Context Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T05:16:19.839437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:16:19.839437Z digest=sha256:d6157156352660f701cdc21171dacf5ec39dda88fac7f9dcb5119f03d76700ad

Observation b66905c9-9f50-4e79-91d7-c9bb567967f5 · inbound

Transformers versus the EM Algorithm in Multi-class Clustering cites this paper.

Transformers versus the EM Algorithm in Multi-class Clustering Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:25.061703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:13:25.061703Z digest=sha256:1826b8ab076f1f163aa14a03ae8032cc6c4f167322251c5fd7befe0ebca303e1

Observation cdcfbd51-fa8a-40ac-97a3-b0915cb0119c · inbound

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias cites this paper.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.990906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.990906Z digest=sha256:ecdc3b0c154beac2443a70d7de18a95a121d6f61e464e637c91997ba0d798d3b

Observation a597f3a7-8913-43cb-bfa4-167279020fab · inbound

Learning Compositional Functions with Transformers from Easy-to-Hard Data cites this paper.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.110434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.110434Z digest=sha256:cf8c314c1a6944f0e1ec0cae302f50bad185b0d274d43844e2f3485c1b787fa6

Observation b0a96b15-0d08-4d63-bc77-d4a28a323726 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:35.674370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:35.674370Z digest=sha256:5336dbff716e029cebc12017bcacb76e7a0567d31bf043bde68a5ef943716021

Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.068890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.068890Z digest=sha256:219937d758ded5b01ce9e6b91b4978398b1edd2c209e211e2e065f185b7b7eb7

Observation 9de3bf8a-670e-4ffd-8d13-05f02c5f9db0 · inbound

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality cites this paper.

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:38.662534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:38.662534Z digest=sha256:5fb1b9861ff4336fcad061eb774949b5b7d02d5ad671ecc05c5d68381f8e4dc5

Observation b67edb81-f98b-49eb-beb4-e52eea0919f0 · inbound

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression cites this paper.

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:17.436794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:10:17.436794Z digest=sha256:8786d489d787d2f5e6b51d480a5b8f305d8db4db59ad12b51234864443e20465

Observation c25d9971-bf1e-4d93-a035-a836d7cc835e · inbound

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention cites this paper.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.033438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.033438Z digest=sha256:011999c9b73f2a422d5d24952749935f9a701c2b7f386bc36f1b9a77f1b8450a

Observation 18964baf-208c-49d5-a14c-5e94b62590f0 · inbound

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge cites this paper.

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:13.762801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:13.762801Z digest=sha256:de034b52adeb9de09141a1028129741002f0ec8b16912a8f9253b28fe02bf41a

Observation 99309c63-be3c-48db-9006-f03ff2e86828 · inbound

How Can Mamba Learn In Context with Outliers and Generalize Provably? cites this paper.

How Can Mamba Learn In Context with Outliers and Generalize Provably? Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:29:29.816756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:29:29.816756Z digest=sha256:a6458786a035d599ee796bc9b7bff7564cd72142200a626f1114daec18d69896

Observation e4e110b9-28e3-49d1-876c-6439086ea6e7 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:09.961879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:09.961879Z digest=sha256:3e80919e668af247bd95dd75c8eab7ab2b0e97d9dd4fcc9db48408b174a9b174

Observation 6b06d115-a9c7-4bdd-a19c-cb2f78b48da0 · inbound

Specialization of softmax attention heads: insights from the high-dimensional single-location model cites this paper.

Specialization of softmax attention heads: insights from the high-dimensional single-location model Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:04:54.999021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:04:54.999021Z digest=sha256:21606bfcd3b8eac7b57f6164df59a6fd73d72d617848a8674edcb8670e462753

Observation 83271fa1-e835-4869-b7bb-96e36e77611f · inbound

Learning to Adapt: In-Context Learning Beyond Stationarity cites this paper.

Learning to Adapt: In-Context Learning Beyond Stationarity Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:59.438462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T16:30:36.771589Z digest=sha256:4951b25b391d3636c710ed0273df2fc024f0a85afe9a1192a4226233cb0c8edc

Observation a096d3db-07f8-4329-b978-d42eb840abcb · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:49.981802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:07c266addfcda01251495f37847a49c05024fe259f51b185976eb9246976ca4e