Pith. sign in

Paper Citation Record · LEDGER

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2505.00926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00926 v3

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:39:55.113761Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:57:10.776717Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:42:49.961290Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cdcfbd51-fa8a-40ac-97a3-b0915cb0119c · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.990906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.990906Z digest=sha256:0d0af4bde1a05afe4c1d4c4157e72c4989faf252030f005fd14e3f639bbb0400

Observation 18158418-9dd5-489e-8a7a-012a1f23fc06 · outbound

This paper cites Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.995663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.995663Z digest=sha256:d62e98a0b4b5ce0395c7c5eef797ede87574c10eb34a93d57c10be34f11778f6

Observation 799d9f0c-6b1f-413f-b6b3-c91d44b9317d · outbound

This paper cites Overcoming a Theoretical Limitation of Self-Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Overcoming a Theoretical Limitation of Self-Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.000654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.000654Z digest=sha256:a01b5b4625e9f13d8459e6cf0b43e9859de3de713882567bc18d9083c5320f92

Observation a7bf4afd-cb6f-4825-8f2a-7f3396c30c9d · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Optimization and Generalization of Multi-head Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.010451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.010451Z digest=sha256:5a61a6897c738a2968b1196a4ec285ca365bf75b65d2c685e7b368477abd7c94

Observation 8f1bf8ba-7431-4872-8a8b-7c2865077a0f · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.014925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.014925Z digest=sha256:c5870c17411910c9fc3584b6aed9c1e0b640052cc792ad0e13d522826945a016

Observation 3232a6ae-4d9f-43c1-aee1-9f9a9fff8c40 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.024520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.024520Z digest=sha256:ebc75e42f1997748df4f73ae209bf28a682b6c8d295170d687eebaafe177da08

Observation b1de9036-333e-45a3-826a-9c832158e0a3 · outbound

This paper cites In-Context Convergence of Transformers.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias In-Context Convergence of Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.033194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.033194Z digest=sha256:4768ac61eb67d2202c740ac49229784153bee19bd6972e2e72d22d549ef4129f

Observation 11777658-631e-4e64-a941-589384b125ef · outbound

This paper cites A Theoretical Analysis of Self-Supervised Learning for Vision Transformers.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias A Theoretical Analysis of Self-Supervised Learning for Vision Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.038985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.038985Z digest=sha256:2c75879df33d7147eb0cd1023e48c0aeb9b9c3b52a99d0ec017104b0f9e3bfb0

Observation 456a1a31-bddc-4d14-ade5-db6d5590bd62 · outbound

This paper cites Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.048091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.048091Z digest=sha256:05c227aecda7f440104e661ccfcaebc2be733028ad5050194e59ca9e85c8aa4d

Observation 06b3ef20-65e2-4861-94a3-5e285381d92a · outbound

This paper cites A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.052535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.052535Z digest=sha256:f92c0e92254db56986db3677be55de4938ec589033b32d043370048d53ce9181

Observation a87cda34-ab35-4ad1-91e6-5dfdb3acde01 · outbound

This paper cites Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.057589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.057589Z digest=sha256:088293aa9373435a20feeb94fbcce12d4051c291fe952ef34f70e43da0bb3b03

Observation b139b00c-2f00-4bc6-aeb4-4bcc0712611d · outbound

This paper cites One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.061423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.061423Z digest=sha256:dd70c70ce2ff9d3f21c60161a66115e493e1257cadf280b777a31022023ee606

Observation 3849a41a-2628-416c-a463-68a4e5478a53 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias The Expressive Power of Transformers with Chain of Thought

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.065880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.065880Z digest=sha256:67577ef4195584d796c6c9c48cb721a32bb64b480e728af1a10e006100ef2056

Observation 579ba2f9-5294-4944-a8ef-4175aa5d9664 · outbound

This paper cites an unresolved cited work.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:39:55.504372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:39:55.070888Z digest=sha256:005debfe4086529f2701320b47fdfba23b233d0a648366064c0ecbb4e79217b0

Observation 27895490-f6d5-404a-ac19-e0b2a9814f1a · outbound

This paper cites Benign Overfitting in Token Selection of Attention Mechanism.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Benign Overfitting in Token Selection of Attention Mechanism

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.078557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.078557Z digest=sha256:ccd49edc31699df471d3fc8b69066d94aeac4d7acf3bd62a35f6bda2bbf6eff7

Observation 439fd820-6ffa-4b94-a19f-4ca19600d8d0 · outbound

This paper cites Implicit Regularization of Gradient Flow on One-Layer Softmax Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Regularization of Gradient Flow on One-Layer Softmax Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.082521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.082521Z digest=sha256:42274d491017e40b9c2e5b597a64e90af59c508e36fa126d3a6ac590e34451b7

Observation 7c15a587-fc9f-4354-9f9f-9ef84f2a3456 · outbound

This paper cites Transformers as Support Vector Machines.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers as Support Vector Machines

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.086987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.086987Z digest=sha256:9d44def74a149371555fa1f168eb9c222e032387acc1e0df9b4dcfc059552a8c

Observation 379d2c9c-b268-487a-aa25-0cdb2c8605b7 · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.092241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.092241Z digest=sha256:31949ca2f740d9d0938fb1c91f0785a570caf5f07d2072d75cb23982a437e1c1

Observation 6a95aef2-5781-4583-a80b-64ac21e2cd4e · outbound

This paper cites Implicit Bias and Fast Convergence Rates for Self-attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Bias and Fast Convergence Rates for Self-attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.096841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.096841Z digest=sha256:55c0b5780f94e7ea60019a1c61d551ebfa1343ee611b7311aec10bf75bcaa45b

Observation 471940bf-60e2-48b6-bc36-428bcb39696a · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.101972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.101972Z digest=sha256:91a238233ece83e5b24842817be8aaec8c1446379a287c9b0faf53242d02cd98

Observation 2aebcc4d-0092-45f9-8b07-40dd5df9ed25 · outbound

This paper cites Trained Transformers Learn Linear Models In-Context.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Trained Transformers Learn Linear Models In-Context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.105878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.105878Z digest=sha256:1c8a68648f8c7f686a9f5b6962cb40005e03cc97d302652db5f4a5d2ec700e23

Observation b0c9f10c-2b93-4a6e-91aa-7cbb25e75fc3 · outbound

This paper cites Auxiliary Lemmas and Equations Lemma A.1 (Gao & Pavel (2017)).

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Auxiliary Lemmas and Equations Lemma A.1 (Gao & Pavel (2017))

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:39:55.490887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:39:55.109869Z digest=sha256:7ad817c7cf39126fad65a82485c12eb7f7a93720b332589493b02adf1a40d6b0

Observation 704aca31-5cd0-4508-b99a-4b474790b8d9 · outbound

This paper cites This can be done by noting that⟨u2,Ew 1 −E2 ℓ⟩≥ Ω(η) forℓ̸=ℓ0.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias This can be done by noting that⟨u2,Ew 1 −E2 ℓ⟩≥ Ω(η) forℓ̸=ℓ0

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:39:55.478138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:39:55.113761Z digest=sha256:d0ba3b426640bb1badcd436b5b1ede262fcfdf07862f9257188ff9c86cda3a1e

Observation 0beb3973-5f06-4848-9285-64e2c0825a93 · outbound

This paper cites Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.019942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.019942Z digest=sha256:3f5ec066c9f672cee5f32484cd3058b5e06e757c4af52300c5471df13a9e6b77

Observation 67f1ecc1-9c23-48f5-b26a-ca3dbb3303aa · outbound

This paper cites How Transformers Learn Causal Structure with Gradient Descent.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias How Transformers Learn Causal Structure with Gradient Descent

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.074860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.074860Z digest=sha256:12e04f0793cb7916950108e6defa9fd12deebaf41e6278bd91e7443e150ff709

Observation dd26e19c-2269-43b7-b139-cc5fea7b213b · outbound

This paper cites Why are Sensitive Functions Hard for Transformers?.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Why are Sensitive Functions Hard for Transformers?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.028824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.028824Z digest=sha256:3ecaa3b0833ac6292fbce0c4f0d3f04c5dc910fac2f98ec47d221f8d8c7007e1

Observation 267c3f93-77d8-43d0-ba85-389b4d88ce17 · outbound

This paper cites Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:39:55.311900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:39:55.043170Z digest=sha256:060ccc6d5cf54a7272e7482bbb80b834e036f02709c6816704dc04feaf72c685

Observation fb8d970e-59c1-4cb6-a089-55a7073a6925 · outbound

This paper cites Superiority of Multi-Head Attention in In-Context Linear Regression.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Superiority of Multi-Head Attention in In-Context Linear Regression

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:39:55.407581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T04:39:55.005255Z digest=sha256:c38c3b9283e7555ff682837a369c182a976fabaa5240a7789158694e95df869a

Observation 6d55574b-e675-4f26-ac77-97799f6e2fa2 · outbound

This paper cites Provably learning a multi-head attention layer.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Provably learning a multi-head attention layer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.985634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.985634Z digest=sha256:3d4734b76d5ed39679e51ab2428edfcbe0b6316795220da48563fd5f993924f6

Observation 8fbe3cd4-788c-44c5-8cde-21f2c77b66ea · outbound

This paper cites On the Ability and Limitations of Transformers to Recognize Formal Languages.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Ability and Limitations of Transformers to Recognize Formal Languages

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.980591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.980591Z digest=sha256:ddc1d0426539c2f4f348eb72d3053513703b6c85630a30cdb46bca663a67a2e3

Pith citing papers

Observation b54f4b11-435e-45a8-9ac1-9f884743d67c · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:10.776717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:10.776717Z digest=sha256:602b75b92744dfd20bccb8b660af7116f74646dfdba9c07206b2fce019263e1b

Observation 16c8f18c-5384-421d-95be-6c837185879a · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:49.962665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:750112d5c4b40d6633fac7b124b345a0a9ae4181be7f20b7b5a917c0c913e1df