Pith. sign in

Paper Citation Record · LEDGER

Two Heads Are Better than One: Simulating Large Transformers with Small Ones

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12220 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:12:48.259338Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:05:58.230449Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b51397b-9819-4a56-bba0-0277ff4eb8cd · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Zoology: Measuring and improving recall in efficient language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.896786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.079998Z digest=sha256:e0697172000f64e4a141979b9024727a67e18d0c6ada20fc2dc75bb7d4cc7492

Observation af1324c6-0fab-4832-a706-76dc132db978 · outbound

This paper cites Fast attention requires bounded entries.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fast attention requires bounded entries

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.880702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.085462Z digest=sha256:b4303c5aca76f3bb19634928c3e120d5238dea7300f76042e54ad6da70da44c0

Observation 8a8047c2-678c-45a3-a1ac-422c5871ea0f · outbound

This paper cites Fundamental limitations on subquadratic alternatives to transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fundamental limitations on subquadratic alternatives to transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.865671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.090223Z digest=sha256:608ee0417959bb3b4405e0c53ca72998eb51e35292f9de94ce923d2ff25d9ede

Observation 7945a37a-e689-418e-b8c0-acba451ad33f · outbound

This paper cites On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.850871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.095765Z digest=sha256:a10aa4162cd410f9389ff63ae787d6faaf1b7e05fbee9f4b5b1ddc6f3180e049

Observation fa65ef42-f46d-40c2-93c3-37b94a062016 · outbound

This paper cites Separations in the representational capabilities of transformers and recurrent architectures.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Separations in the representational capabilities of transformers and recurrent architectures

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.835057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.100144Z digest=sha256:fda4262e423a303794c6253b59ac1f358c3f73f7cf7097eb6fc3a2de91f0a4b4

Observation 024fca3b-7758-4ed9-9ec7-c685edee9285 · outbound

This paper cites Longformer: The Long-Document Transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.104615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.104615Z digest=sha256:a98fd9ffcd17cc49c46851fd6bbabfbdd9397735cd835dbbe0c7156498f88ea6

Observation addbc6b7-327f-4df2-a6d5-7a5383a7a3a4 · outbound

This paper cites An exploration of hierarchical attention transformers for efficient long document classification, 2022.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An exploration of hierarchical attention transformers for efficient long document classification, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.819050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.110449Z digest=sha256:8995ab50ebf2ffb078f567dc811673823b2fd5256adf467d994a0422c7ea0157

Observation d63ca6e3-07d2-4db9-ada5-861a5bfe1370 · outbound

This paper cites Colwell, and Adrian Weller.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Colwell, and Adrian Weller

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.114975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.114975Z digest=sha256:da250ee2dc0081bf74ef2cb3c3dc80ab12a891de404992d725f62ed376db53ca

Observation 208bc6d9-2b45-4ef8-adae-3785c449204f · outbound

This paper cites End-to-end object detection with transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones End-to-end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.794430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.119593Z digest=sha256:23b691202c8782cd1b78298118436daf560d7e91d7b2c047dfd435ddba452ddc

Observation c0d9153c-e51b-4936-abef-f0a471cc6c8c · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.123953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.123953Z digest=sha256:e23eae448d8e8ee4a067ee0b09b1cf3399386d63672ea429e7528983839397b3

Observation 74419346-96a3-4ec1-940e-a64dd8201936 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.779890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.128702Z digest=sha256:734ba690b81423bb4cc76fc2ad4e5cba358283f090e42bfaa422a5d9ad8f189e

Observation 02e50453-d976-4a8e-a032-008066dfdcaf · outbound

This paper cites Etched is Making the Biggest Bet in AI.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Etched is Making the Biggest Bet in AI

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.765080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.133124Z digest=sha256:9e5a95e945c2c97c685f43e09202e0229c515c89e658a33eb0877a7609d5f18d

Observation f4cb5a0f-b75d-436c-996d-6573336e49c3 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.138043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.138043Z digest=sha256:01369f8d986fc5146a0d6a924f58d3eca79399a00ef292cbe6d1cdcdf0a3c82c

Observation a02b2ac7-3259-4954-8b3e-8da43426ddb9 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Theoretical limitations of self-attention in neural sequence models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.748865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.143413Z digest=sha256:c4b9bf14ab1d2b8878cfe6921b6b09684c3dc7428229ff22f3a2e2dda0081058

Observation 38ea6915-586e-4ae5-941a-ada67e86cc74 · outbound

This paper cites Hyperattention: Long-context attention in near-linear time.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hyperattention: Long-context attention in near-linear time

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.734346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.148089Z digest=sha256:6f0ed6910f6436b6213edafb560a4bd55f3c817ec98f1482a888b01715a32f70

Observation 08d07ec7-a333-4c40-8c90-2aa45499d470 · outbound

This paper cites Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.719060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.152253Z digest=sha256:15f25cd424dfacb2735b57a78021ede2234e917d4a6f113f8c9ca33e16fda04d

Observation fe1866d1-3cbd-4e71-9b22-14369cd71c7c · outbound

This paper cites Multilayer feedforward networks are universal approximators.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Multilayer feedforward networks are universal approximators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.703811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.156366Z digest=sha256:5b831d1f6447264e49a52337738c0fc5a767200da133951c093ef8ce966f174f

Observation 38e00788-d632-489c-995f-418982c0cfe1 · outbound

This paper cites On statistical rates and provably efficient criteria of latent diffusion transformers (dits).

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On statistical rates and provably efficient criteria of latent diffusion transformers (dits)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.689004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.160515Z digest=sha256:be07f27dcee37f3d4bea17f2847cdcaa8f810de3c5ceaaa2188531df76293bdd

Observation 7cdac334-4f01-435b-8d57-51f8908b4925 · outbound

This paper cites Kakade, and Eran Malach.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Kakade, and Eran Malach

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.165303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.165303Z digest=sha256:724ff403461d4b244579067a071838309cba1692df2c39007ef30734159699e8

Observation acebf52c-beb3-4a63-ab6c-4ac01dd9cf27 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An image is worth 16x16 words: Transformers for image recognition at scale

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.663817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.170867Z digest=sha256:4a810b7404ce03a49882a2c3bba377af77790e1fecfa9123c7e16d688c379789

Observation e9d094b8-0055-414d-9fb4-a96198d7feb0 · outbound

This paper cites RNGD preview: The world's most efficient AI chip for LLM inference.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNGD preview: The world's most efficient AI chip for LLM inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.650041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.175070Z digest=sha256:f883255159f3fd3c5af3314146d43a8aada9f4c8f6488335ccd042fdf18e44f9

Observation d80f281d-47c8-4a70-a344-5a40cd0df65d · outbound

This paper cites Reformer: The efficient transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reformer: The efficient transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.635829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.178891Z digest=sha256:ae1fb45eb3807fdfdbf41c752ebade5a3b6efd620fe069206e8191d597bb0410

Observation 278afdd6-e958-4ee1-906d-af3f472e816a · outbound

This paper cites Polysketchformer: fast transformers via sketching polynomial kernels.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Polysketchformer: fast transformers via sketching polynomial kernels

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.622136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.182787Z digest=sha256:1080bfecf13fd2c027389baca4b477462ca09efb0d7edbcb758c0cd19f0328e2

Observation 48ed1d42-b864-410a-bf0b-22644b5e50be · outbound

This paper cites Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.608236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.186811Z digest=sha256:85358c9ba24923c3d2415ca61a2b3050d960772b492a06aaf7bc93295c943504

Observation ed2960ec-a1f8-4d69-88f3-f2648a08b744 · outbound

This paper cites On the expressive flexibility of self-attention matrices.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the expressive flexibility of self-attention matrices

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.593022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.190765Z digest=sha256:a667cf6450cdf3596895519ef359deaa8d475ee0932fe1a06b5303edc1fd3807

Observation c29f4383-d564-47e4-b9a4-2b207dacee37 · outbound

This paper cites Hierarchical transformers for multi-document summarization.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for multi-document summarization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.578990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.194683Z digest=sha256:df28f079064076df35be57dc384ac8ec1d50394d5bf0d803baafc4c350d7d0c5

Observation d90d4e61-9a47-4870-896a-c408cef07b8b · outbound

This paper cites The parallelism tradeoff: Limitations of log-precision transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones The parallelism tradeoff: Limitations of log-precision transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.564989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.198545Z digest=sha256:8f91335a8a43dff37f1376c8ed267978ed5941d93dc9fa994caecf3154923093

Observation fc1f91c6-6d9a-44ea-9d52-57cf35de5df0 · outbound

This paper cites an unresolved cited work.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:12:48.550241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.202818Z digest=sha256:58d95da7e0d3e0b3893d187f767417d5e83f86c169f0671f752bcb3ddacd852d

Observation 3bb1e02f-d1fa-4f28-9bbd-3aa12a639221 · outbound

This paper cites Language models are few-shot learners.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Language models are few-shot learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.535978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.206861Z digest=sha256:25f056ee412a849182d1f2f8b6629a9f129583853b9827996b7c78db01d0f406

Observation 3b258ad8-d2ad-446f-badb-12896773279a · outbound

This paper cites Hierarchical transformers for long document classification.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for long document classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.521713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.210920Z digest=sha256:c4581d8105fe37a1151beadcef610607d1aa2c4004023cacd559417557d62b63

Observation b661e8a2-a0a8-4843-add4-4105573376c4 · outbound

This paper cites Representational strengths and limitations of transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Representational strengths and limitations of transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.507820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.214838Z digest=sha256:1a85aa0871daea81d7274386af23955007aac5a5461c5f812d11c6d8fa2ccd6b

Observation 3ef4a2f6-b742-4a37-a5c6-ff8c5eba978d · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Transformers, parallel computation, and logarithmic depth

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.493360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.218877Z digest=sha256:2b7f57305d24426585fa675ebfda3220d2914df6dba2fcbb30951314d23e0eaa

Observation 44e1558a-3115-4c57-9557-a69b1b43a757 · outbound

This paper cites What formal languages can transformers express? a survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones What formal languages can transformers express? a survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.477544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.223004Z digest=sha256:2702a090c9197b0d8adbba5535a2384a947aa7c2fdb56e08ababce56bddbc2a2

Observation c1d55e78-58a6-4b59-9ea1-a56d3a76e0b8 · outbound

This paper cites Efficient transformers: A survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient transformers: A survey

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.462466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.227406Z digest=sha256:e0ae4f0bd49756cceb0ea0f8fe4111178c95f8d26e94730af9ceb40da2856c64

Observation 9445936f-2977-44c6-af2d-d60cdfd6cd04 · outbound

This paper cites Schmidt, and Stephan Peitz.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Schmidt, and Stephan Peitz

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.446971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.231424Z digest=sha256:7ec155c874b97a56c61e17ccdea222c794af600ed258d260d20c75efb354faf0

Observation b3536e4d-94ec-4bcd-bfec-8e67154f4f8e · outbound

This paper cites Gomez, ukasz Kaiser, and Illia Polosukhin.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Gomez, ukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.429843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.235774Z digest=sha256:f090d60dcd369c67a12ee968a1fb569abf409920bed9819010d0a047d5d5e95c

Observation 80d2ebdf-30d9-43e7-b2e9-9935b9e57e94 · outbound

This paper cites RNN s are not transformers (yet): The key bottleneck on in-context retrieval.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNN s are not transformers (yet): The key bottleneck on in-context retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.414278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.240370Z digest=sha256:995cc6f96b48656535a50aef9b955151566f30d86d9d23c4db0e642127d167d0

Observation de93e740-306e-4720-b64c-b46d5a14bddd · outbound

This paper cites LightSeq2: Accelerated Training for Transformer-based Models on GPUs.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones LightSeq2: Accelerated Training for Transformer-based Models on GPUs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T01:12:48.303846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.245834Z digest=sha256:23275546fab74d42c0e162d0f80eafdcd6e189be1555b35e1556e2733ba82762

Observation 7d215dfb-3d60-43fe-8962-399d38217616 · outbound

This paper cites Efficient streaming language models with attention sinks.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient streaming language models with attention sinks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.398365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.251148Z digest=sha256:fc8b956ab860f852825d9dedf8e9f162088556dea86c2dddc95d1fa118722f95

Observation 12d921ac-cc04-4481-b458-297b95299f53 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.382813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.255183Z digest=sha256:a6e473be874b923ccf3ae77c4d8cffc02a228dabc48e2818acfe8e61e87393bd

Observation c1d80b37-089f-44c1-868f-5cd1ac530dc0 · outbound

This paper cites Reddi, and Sanjiv Kumar.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reddi, and Sanjiv Kumar

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.366538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.259338Z digest=sha256:a9228abc3794fd2861b6a74949617fab713fc88b09e382593e959b31cfbb187d

Pith citing papers

Observation 795242bf-3d0f-49a8-96ad-9f827583f903 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Two Heads Are Better than One: Simulating Large Transformers with Small Ones

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.232597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:71833441d2d33927922adb341c7bf2d68c2f2a45926801da479059a58abd2696