Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2402.19442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19442 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:34.110434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:42:49.980406Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a597f3a7-8913-43cb-bfa4-167279020fab · inbound

Learning Compositional Functions with Transformers from Easy-to-Hard Data cites this paper.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:34.110434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:34.110434Z digest=sha256:39124aa0083c0d9ddd557ab52c70f54d4b6d658d54eb06faf5bfa1864a560497

Observation b0a96b15-0d08-4d63-bc77-d4a28a323726 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:35.674370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:35.674370Z digest=sha256:a9ce981b33cd6bc01e1826b2ef4dc61b3a6243d0160a0552e68f8b11f198dfe1

Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.068890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.068890Z digest=sha256:25c0dde1207e9ddb64776638edeec53f394e7a8dfd8325a79cfcdbb7370d791d

Observation 9de3bf8a-670e-4ffd-8d13-05f02c5f9db0 · inbound

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality cites this paper.

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:38.662534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:38.662534Z digest=sha256:4097ffe78a13854198fc5573e92b0a35fc9513506f59448ee44eaf03965f97de

Observation b67edb81-f98b-49eb-beb4-e52eea0919f0 · inbound

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression cites this paper.

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:17.436794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:10:17.436794Z digest=sha256:91690e856e19d512c3d63caa8ec880d83bf18e3f229a4d46b7aaa54c7cec5394

Observation c25d9971-bf1e-4d93-a035-a836d7cc835e · inbound

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention cites this paper.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.033438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.033438Z digest=sha256:b82999dbfdfcd1342fbb836b7319830dd941609e4bdeb54810c24b6cca0756d2

Observation 18964baf-208c-49d5-a14c-5e94b62590f0 · inbound

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge cites this paper.

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:13.762801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:13.762801Z digest=sha256:7466d1edf2d1188cc686b7efa01558a6df50069c652e96002f4034b6763719d3

Observation 99309c63-be3c-48db-9006-f03ff2e86828 · inbound

How Can Mamba Learn In Context with Outliers and Generalize Provably? cites this paper.

How Can Mamba Learn In Context with Outliers and Generalize Provably? Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:29:29.816756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:29:29.816756Z digest=sha256:094e29d485da6b64a21e90849820c9984f2951f8ea1f69c7e0866488f77d00e7

Observation e4e110b9-28e3-49d1-876c-6439086ea6e7 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:09.961879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:09.961879Z digest=sha256:6c1f8961d671c191664730e0aca9d96791a317e286a271bdda7ac9d2cf3170d4

Observation 6b06d115-a9c7-4bdd-a19c-cb2f78b48da0 · inbound

Specialization of softmax attention heads: insights from the high-dimensional single-location model cites this paper.

Specialization of softmax attention heads: insights from the high-dimensional single-location model Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:04:54.999021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:04:54.999021Z digest=sha256:ae41b3f31625c3d6d83dfa2c325218071a7a9a6ddfa1c7d2616f10fa2a9daecf

Observation 83271fa1-e835-4869-b7bb-96e36e77611f · inbound

Learning to Adapt: In-Context Learning Beyond Stationarity cites this paper.

Learning to Adapt: In-Context Learning Beyond Stationarity Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:59.438462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:30:36.771589Z digest=sha256:eb3b2ce2cb116d3c25ff3eab9712c4cc619f05e0fd703ce261acbec8fdf7768c

Observation a096d3db-07f8-4329-b978-d42eb840abcb · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:49.981802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:459453c18102dc7ac22b98a52934342ea661f63dbcf9ddf51a1d5a7fcab0c1e5