Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:12:48.259338Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12220.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:12:48.259338Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T09:05:58.230449Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b51397b-9819-4a56-bba0-0277ff4eb8cd · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Zoology: Measuring and improving recall in efficient language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation af1324c6-0fab-4832-a706-76dc132db978 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fast attention requires bounded entries
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a8047c2-678c-45a3-a1ac-422c5871ea0f · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fundamental limitations on subquadratic alternatives to transformers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7945a37a-e689-418e-b8c0-acba451ad33f · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fa65ef42-f46d-40c2-93c3-37b94a062016 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Separations in the representational capabilities of transformers and recurrent architectures
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 024fca3b-7758-4ed9-9ec7-c685edee9285 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Longformer: The Long-Document Transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation addbc6b7-327f-4df2-a6d5-7a5383a7a3a4 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones An exploration of hierarchical attention transformers for efficient long document classification, 2022
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d63ca6e3-07d2-4db9-ada5-861a5bfe1370 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Colwell, and Adrian Weller
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208bc6d9-2b45-4ef8-adae-3785c449204f · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones End-to-end object detection with transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c0d9153c-e51b-4936-abef-f0a471cc6c8c · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74419346-96a3-4ec1-940e-a64dd8201936 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 02e50453-d976-4a8e-a032-008066dfdcaf · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Etched is Making the Biggest Bet in AI
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f4cb5a0f-b75d-436c-996d-6573336e49c3 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a02b2ac7-3259-4954-8b3e-8da43426ddb9 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Theoretical limitations of self-attention in neural sequence models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 38ea6915-586e-4ae5-941a-ada67e86cc74 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hyperattention: Long-context attention in near-linear time
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 08d07ec7-a333-4c40-8c90-2aa45499d470 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe1866d1-3cbd-4e71-9b22-14369cd71c7c · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Multilayer feedforward networks are universal approximators
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 38e00788-d632-489c-995f-418982c0cfe1 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones On statistical rates and provably efficient criteria of latent diffusion transformers (dits)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cdac334-4f01-435b-8d57-51f8908b4925 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Kakade, and Eran Malach
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acebf52c-beb3-4a63-ab6c-4ac01dd9cf27 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones An image is worth 16x16 words: Transformers for image recognition at scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e9d094b8-0055-414d-9fb4-a96198d7feb0 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNGD preview: The world's most efficient AI chip for LLM inference
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d80f281d-47c8-4a70-a344-5a40cd0df65d · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reformer: The efficient transformer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 278afdd6-e958-4ee1-906d-af3f472e816a · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Polysketchformer: fast transformers via sketching polynomial kernels
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48ed1d42-b864-410a-bf0b-22644b5e50be · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ed2960ec-a1f8-4d69-88f3-f2648a08b744 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the expressive flexibility of self-attention matrices
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c29f4383-d564-47e4-b9a4-2b207dacee37 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for multi-document summarization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d90d4e61-9a47-4870-896a-c408cef07b8b · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones The parallelism tradeoff: Limitations of log-precision transformers
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fc1f91c6-6d9a-44ea-9d52-57cf35de5df0 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3bb1e02f-d1fa-4f28-9bbd-3aa12a639221 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Language models are few-shot learners
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3b258ad8-d2ad-446f-badb-12896773279a · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for long document classification
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b661e8a2-a0a8-4843-add4-4105573376c4 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Representational strengths and limitations of transformers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3ef4a2f6-b742-4a37-a5c6-ff8c5eba978d · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Transformers, parallel computation, and logarithmic depth
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 44e1558a-3115-4c57-9557-a69b1b43a757 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones What formal languages can transformers express? a survey
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c1d55e78-58a6-4b59-9ea1-a56d3a76e0b8 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient transformers: A survey
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9445936f-2977-44c6-af2d-d60cdfd6cd04 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Schmidt, and Stephan Peitz
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b3536e4d-94ec-4bcd-bfec-8e67154f4f8e · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Gomez, ukasz Kaiser, and Illia Polosukhin
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80d2ebdf-30d9-43e7-b2e9-9935b9e57e94 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNN s are not transformers (yet): The key bottleneck on in-context retrieval
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation de93e740-306e-4720-b64c-b46d5a14bddd · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones LightSeq2: Accelerated Training for Transformer-based Models on GPUs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7d215dfb-3d60-43fe-8962-399d38217616 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient streaming language models with attention sinks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 12d921ac-cc04-4481-b458-297b95299f53 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c1d80b37-089f-44c1-868f-5cd1ac530dc0 · outbound
Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reddi, and Sanjiv Kumar
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 795242bf-3d0f-49a8-96ad-9f827583f903 · inbound
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Two Heads Are Better than One: Simulating Large Transformers with Small Ones
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.