Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:18.403313Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2608.06776.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:18.403313Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:48.968402Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-14T04:15:49.290160Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 05db1fb3-1915-4d4d-8c9c-42167c5cf6af · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f04f26b-4471-4ce9-bd8c-9554a47a806a · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Self-attention networks localize when qk-eigenspectrum concentrates
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cd6eb2e9-caea-43e5-ad11-61852b497662 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Birth of a transformer: a memory viewpoint
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 68441002-f3d1-488f-844c-f3690651deba · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b2cc58-65ff-4a71-be7b-5feffc94c991 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Universal Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 067bf91f-cb3a-40d2-ab3e-795be6d9a3c1 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the optimization and generalization of multi-head attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e9cbbafc-1209-481a-8049-06b45c4e698f · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Optimization and Generalization of Multi-head Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d32f044-0fca-489c-87a3-948b4f932167 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e466f636-4ef9-494b-b147-0b4bf080a84d · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models S., Hu, W., and Lee, J
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cb061834-8ad3-4064-af60-2b39a56c5cab · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models A mathematical framework for transformer circuits
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943218e2-4634-4659-bb85-c76682efc00d · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models The emergence of clusters in self-attention dynamics
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d979015-6b34-4794-b445-e49d6f16c970 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Wallace, B
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aeb751e1-e6a1-464f-8c3e-9554802a074b · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Clustering in causal attention masking
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fbb1ea4c-32ed-400f-b5bf-9c2a5f4e0b06 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Transformers in Speech Processing: A Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed667f7c-1045-4aa7-aeb8-e55d0dee41c1 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models How do transformers learn topic structure: towards a mechanistic understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98b45860-5b3b-499c-bdfd-1ab759a7fca4 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Mechanics of Next Token Prediction with Self-Attention
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b8c52ce-c7a1-41a4-aff1-6afe9c5d5b1b · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the dynamics of training attention models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 075fc0dd-a114-433e-9230-450eb2083918 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models M., Biemann, C., Goyal, P., and Mukherjee, A
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 51cf7cdf-3d68-464b-8078-1685cf2ed3df · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models In-context learning and induction heads
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b99241-6a03-4d27-87eb-66ac25ca15bc · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models N., Vashisht, R., and Ramaswamy, H
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5fe12347-73ea-4936-9372-1998bc77a2d9 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is turing complete
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b8e5479-3e76-4b9b-8c91-81e0fd0b96e3 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Improving language understanding by generative pre-training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be5a6b6-9d0b-4cd2-9274-9fba9d662ea9 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Scan and snap: understanding training dynamics and token composition in 1-layer transformer
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ffb953c-5610-48a1-adee-fdafc300c806 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Ramaswamy, H
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7f632b6b-0afa-41fb-9f36-6a46896f5417 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models N., Kaiser, L
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b029877b-2e9f-4876-aa2c-e7bad65ca26b · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Are Transformers universal approximators of sequence-to-sequence functions?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85017a99-ae25-4307-a77e-cbcdae1831a7 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is All you Need , url =
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9962a82-cb5f-4b3b-bfb0-52dc99b649be · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2ce633c1-dec7-4c60-afb2-4cd12f19af38 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Kumar, Sanjiv , biburl =
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5d45f57c-fe0f-49be-8482-35a94a84243d · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Computational Power of Transformers and Its Implications in Sequence Modeling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4cb7314-b594-4f33-8148-cd3c81dac9e1 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8377c4-f28d-453a-acb9-f3fd5d7cd7fd · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d77dc4ac-07c0-43b2-ab88-db6a58df865a · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is turing complete , year =
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 61fbacbd-5c02-4471-afc2-08d979fb5a98 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Ba, Jimmy , biburl =
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c4106af3-60f1-4649-bdb5-0c2563327275 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b29081aa-6747-4c23-881c-16e2651e7f3f · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2021 , journal=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08a058c-d101-4628-a59b-99d93c1190fb · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models In-context Learning and Induction Heads
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2eb62838-e40d-4716-ab1a-13ba5b48a26f · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Birth of a transformer: a memory viewpoint , year =
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9be6d413-dccf-4002-8194-ad25097c2cab · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models SQ u AD : 100,000+ Questions for Machine Comprehension of Text
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 068cf1d7-7f45-4414-a4b1-7db80a11e364 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Assessing the Ability of LSTM s to Learn Syntax-Sensitive Dependencies
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f7765e-f0c2-4d95-b874-b2dd21a4be76 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models AAAI Conference on Artificial Intelligence , year=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 23c39edd-bddc-4ee3-a11c-98918278029e · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9c8e705f-8870-49fc-9ad2-ebe22dc20016 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , url =
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8d2a32-6d35-4dcc-a34c-982f5f57ddba · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Language Models are Few-Shot Learners
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7b9722-6608-46af-8395-362f5039b842 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45d4986-e646-41c3-bde9-6c52bac9508d · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b585f6d-0d2c-489a-8b1f-e72b224ff1f3 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aa5540ea-bd45-4b08-9017-ac9559d068d4 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Hu, Wei and Lee, Jason D
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f437a744-300c-42be-bb1a-ee39277e3456 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2022 , journal=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7163e939-6afe-489e-9b05-27f555586377 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d1602f52-1aff-4bea-b548-b596476f0b21 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2025 , eprint=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2296ce4-2ff9-4269-9729-6c34e0085c92 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is not not Explanation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 521285b6-e782-406a-bf6d-d1f40b979d6a · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models North American Chapter of the Association for Computational Linguistics , year=
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 07c2fcb9-a389-46bd-834e-4e59d86f645a · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Is Attention Interpretable?
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300fb36e-0524-4fe8-af54-54b6a7c665d3 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2024 , eprint=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 68feb1c8-a306-497d-a414-4a0f00296363 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 381831d6-f1f6-428a-a8e6-5cfe10bfe1cf · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 41st International Conference on Machine Learning , articleno =
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 18b0ed52-e575-4366-87cd-110ff5be9c6c · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Transactions on Machine Learning Research , issn=
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 15745298-06e3-45ef-afa0-a94c60961d3f · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence , articleno =
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 44b4543b-df80-4580-911e-9d83d571ef11 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models International Conference on Learning Representations , year=
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e226e27-f37b-4687-8bf0-40f6d8272bc6 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c9a9e8f5-7966-4a6f-992f-1ef36ca5cdb3 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b8a100d-98f6-4ca0-a665-cff24368d6fa · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation af580012-a011-4851-bec7-bd87cb694164 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 40th International Conference on Machine Learning , articleno =
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cd807fa5-061e-4dba-b80b-2eb806503025 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Quantifying Attention Flow in Transformers
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b166557b-8498-4d35-bc1e-b76757058ca0 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models ERASER : A Benchmark to Evaluate Rationalized NLP Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0c266e-9a1d-4525-b9c9-abda41873ce3 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of The 14th Asian Conference on Machine Learning , pages =
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 45bb5563-3142-4a07-b98a-769964221409 · outbound
Faster Query-Key Learning Sharpens Attention in Self-Attention Models European Conference on Artificial Intelligence , year=
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f6db4ff7-a0e7-480d-a701-6470ca84d468 · inbound
Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference Faster Query-Key Learning Sharpens Attention in Self-Attention Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.