Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:39:55.113761Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2505.00926.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:39:55.113761Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:57:10.776717Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T23:42:49.961290Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cdcfbd51-fa8a-40ac-97a3-b0915cb0119c · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18158418-9dd5-489e-8a7a-012a1f23fc06 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799d9f0c-6b1f-413f-b6b3-c91d44b9317d · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Overcoming a Theoretical Limitation of Self-Attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7bf4afd-cb6f-4825-8f2a-7f3396c30c9d · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Optimization and Generalization of Multi-head Attention
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f1bf8ba-7431-4872-8a8b-7c2865077a0f · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3232a6ae-4d9f-43c1-aee1-9f9a9fff8c40 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1de9036-333e-45a3-826a-9c832158e0a3 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias In-Context Convergence of Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11777658-631e-4e64-a941-589384b125ef · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456a1a31-bddc-4d14-ade5-db6d5590bd62 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b3ef20-65e2-4861-94a3-5e285381d92a · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87cda34-ab35-4ad1-91e6-5dfdb3acde01 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b139b00c-2f00-4bc6-aeb4-4bcc0712611d · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3849a41a-2628-416c-a463-68a4e5478a53 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias The Expressive Power of Transformers with Chain of Thought
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579ba2f9-5294-4944-a8ef-4175aa5d9664 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 27895490-f6d5-404a-ac19-e0b2a9814f1a · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Benign Overfitting in Token Selection of Attention Mechanism
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439fd820-6ffa-4b94-a19f-4ca19600d8d0 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c15a587-fc9f-4354-9f9f-9ef84f2a3456 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers as Support Vector Machines
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379d2c9c-b268-487a-aa25-0cdb2c8605b7 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a95aef2-5781-4583-a80b-64ac21e2cd4e · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Bias and Fast Convergence Rates for Self-attention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 471940bf-60e2-48b6-bc36-428bcb39696a · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aebcc4d-0092-45f9-8b07-40dd5df9ed25 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Trained Transformers Learn Linear Models In-Context
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c9f10c-2b93-4a6e-91aa-7cbb25e75fc3 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Auxiliary Lemmas and Equations Lemma A.1 (Gao & Pavel (2017))
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 704aca31-5cd0-4508-b99a-4b474790b8d9 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias This can be done by noting that⟨u2,Ew 1 −E2 ℓ⟩≥ Ω(η) forℓ̸=ℓ0
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0beb3973-5f06-4848-9285-64e2c0825a93 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f1ecc1-9c23-48f5-b26a-ca3dbb3303aa · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias How Transformers Learn Causal Structure with Gradient Descent
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd26e19c-2269-43b7-b139-cc5fea7b213b · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Why are Sensitive Functions Hard for Transformers?
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267c3f93-77d8-43d0-ba85-389b4d88ce17 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fb8d970e-59c1-4cb6-a089-55a7073a6925 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Superiority of Multi-Head Attention in In-Context Linear Regression
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d55574b-e675-4f26-ac77-97799f6e2fa2 · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Provably learning a multi-head attention layer
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbe3cd4-788c-44c5-8cde-21f2c77b66ea · outbound
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Ability and Limitations of Transformers to Recognize Formal Languages
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b54f4b11-435e-45a8-9ac1-9f884743d67c · inbound
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c8f18c-5384-421d-95be-6c837185879a · inbound
Agentic Transformers Provably Learn to Search via Reinforcement Learning How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.