Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T23:05:56.687644Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 3 inbound Pith citation observations for arXiv:2508.00901.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T23:05:56.687644Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:55:10.121913Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-01T17:35:51.333328Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5863d64c-f89d-41e6-89d1-13d758384f0f · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Transformers learn to implement preconditioned gradient descent for in-context learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9263cb36-40c2-47f6-b769-3111ece9d798 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b250b3e-629d-4e98-a566-dbccacca7cfe · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 52ce4e83-416e-4755-9181-98348147bcb4 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Birth of a transformer: A memory viewpoint
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 84d2161e-a267-429d-8f19-ece9b342cb19 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Language models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e3e4f27d-6cd1-488b-ac03-dc1803f1a87a · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Scaling Laws for Associative Memories
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 062bd6b0-43a0-4f38-a1a9-a41b7d86e4b3 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Learning Associative Memories with Gradient Descent
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc3e7a8e-b5b9-4dd3-8e8e-5a2025c77776 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Benign overfitting in two-layer convolutional neural networks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ff181a0-5d14-4ec2-a864-d3ac0d70de10 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Towards understanding the mixture-of-experts layer in deep learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1a7b7cd7-c4d6-4a78-9f4e-10c11430db8b · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82df89cf-70ca-4945-b7ec-772d03367ac6 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Transformer Feed-Forward Layers Are Key-Value Memories
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c337a012-9ada-484b-bc02-887e77929c3e · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Understanding Finetuning for Factual Knowledge Extraction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c54b138c-578a-4ea6-9a19-c6b325a07410 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c79b9456-d1d4-4a60-80d9-326a67f33907 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers In-Context Convergence of Transformers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 17e089b1-a4a6-4f41-adeb-e1763392258c · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Vision transformers provably learn spatial structure
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 693b6d58-a4b3-4cb6-aa09-c9b8cd74746d · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unveil benign overfitting for transformer in vision: Training dynamics, convergence, and generalization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7cc9718f-71dd-49ee-89bb-c509c88c97f0 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Optimal memorization capacity of transformers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8a33e6e2-c8c1-41bb-a24d-1498eb4cb55e · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Large language models struggle to learn long-tail knowledge
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation efcd39cd-ba49-4196-8bd6-e7193b11aac7 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Provable memorization capacity of transformers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d6a51fe-209e-450a-8266-d5203507242a · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Benign overfitting in two-layer relu convolutional neural networks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7a0fcf5-efa5-40d6-a60d-f92613bee5aa · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Deduplicating Training Data Makes Language Models Better
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd2cff31-1120-40a7-b9ec-f60124cdf36e · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Next-token prediction capacity: general upper bounds and a lower bound for transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98571dc9-753e-4717-8ae1-ea886bd49bf9 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6bdd94b1-7785-4547-9142-481f60980056 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Memorization Capacity of Multi-Head Attention in Transformers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ed10931f-24e5-4a47-b46f-53eed430d8a9 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2b5b6738-1257-4614-84fb-2579408cb3cf · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Locating and editing factual associations in gpt
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f83e290-35b7-4ccd-80a0-69c93bc572ee · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers How Transformers Learn Causal Structure with Gradient Descent
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f6500055-9d8b-49dc-8dcf-2c18cfa67d4d · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Understanding Factual Recall in Transformers via Associative Memories
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e3872ae8-077c-42e0-a6ef-ff13c614f70c · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Language models are unsupervised multitask learners
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation db597d97-b182-46ec-9c2e-1f305677ac0e · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Data augmentation as feature manipu- lation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2280d1c-a3a4-42d7-a954-3856dbd84e58 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Scan and snap: Understand- ing training dynamics and token composition in 1-layer transformer
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 59d62af1-a3cd-4273-bb39-e1f1000f3757 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9de1d0b7-1632-483d-a5e2-0f27b96f5bb4 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Rethinking benign overfitting in two-layer neural networks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d245bf62-821a-4499-8cb7-f001155fdc56 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Knowledge circuits in pretrained transformers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98967ec6-cae9-4447-a2cc-23f063076a39 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Are transformers universal approximators of sequence-to-sequence functions? International Conference on Learning Representations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf517ccc-3fbb-4f49-9f22-7fd04b8cbda4 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Trained transformers learn linear models in-context
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b114abe6-ebd9-491d-8722-854deeab8a7a · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Towards a theoretical understanding of the’reversal curse’via training dynamics
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d86208f-6a23-4cea-b185-d9698c641861 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers The benefits of mixup for feature learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9275d688-f560-42b3-a53c-f2368f50236b · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d3d1eaeb-e6a3-4c63-8fb3-99ab63d3c1c6 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 295216a7-5bf3-460f-b129-6ce56fce8e39 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a625a410-8c41-4c57-bdc4-4e2ae50bf63e · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Based on the sentence set T , we construct the pre-training dataset as follows: For each (sj, ri, aj) ∈ T
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98e94caf-d04e-4d28-abd9-8121290fc499 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers We construct the fine-tuning dataset and the test dataset as For each (sj, ri, aj) ∈ T
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fba0194a-03a7-47c2-9c64-8df93d491ef8 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbbb93a2-a461-48f3-9f98-37b7fea1d8a7 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers We assume all token embeddings— sj, ri, aj, for all j ∈ [N] and i ∈ R are generated from random Gaussian distributions N (0, σ2I)
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ee853f0d-b7d9-4e61-b8b9-7532c6fe971b · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers nX i=1 I(j ∈ A i) ≥ K 2q n # ≤ exp − nK2 2q2 . (44) 25 Proof. In each time i ∈ [n], a number j ∈ [q] has probability K/q to selected. By Hoeffding’s inequality, we have: P
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7aa70d31-1560-4218-bea5-e5de6f120a3d · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e2e6550-6ab9-478a-9902-e9fcbb57cf09 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 55e05847-3ee1-459b-8885-5013041cf37f · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 583ff63e-d385-4041-aa99-d694cde1589f · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers (70) Proof
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98690a49-de06-4405-9f29-1755c3d1739b · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers After 2 iterations, we have 1 m mX k=1 ⟨w(T1) I(aj),k, Ξ(T1)([o s j ri])⟩ = O λη1dKs(j) nm
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 555e56af-0ffe-4acf-8c6a-d9485ad77dfb · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fd531e6-32e7-4b0a-ab5c-69c3a89e4746 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 61a7c4a0-89ef-441d-ae30-9c6c76589b08 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd73b812-2074-42fe-857c-c753a89ae53f · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9de4e054-075a-4007-952f-193140735582 · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4f5e8f00-9d7e-4d16-9b23-7fe38d5068fd · outbound
Provable Knowledge Acquisition and Extraction in One-Layer Transformers dX j=1 σjujv⊤ j # i , [RA1(G(tf ))]i = σ1u1v⊤ 1 i . (216) Taking the L2 norms, we have [G(tf )]i 2 2 =
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fca444f4-1977-4d43-8344-ba420ef3a8f3 · inbound
Procedural Pretraining: Warming Up Language Models with Abstract Data Provable Knowledge Acquisition and Extraction in One-Layer Transformers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193e0fc8-0d05-4f64-bd4c-868e6ee334f4 · inbound
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Provable Knowledge Acquisition and Extraction in One-Layer Transformers
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 64ea49a1-5f14-4feb-9e8f-c4367f115cea · inbound
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Provable Knowledge Acquisition and Extraction in One-Layer Transformers
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.