Pith. sign in

Paper Citation Record · LEDGER

Two Heads Are Better than One: Simulating Large Transformers with Small Ones

As of 14 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12220 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:12:48.259338Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:05:58.230449Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b51397b-9819-4a56-bba0-0277ff4eb8cd · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Zoology: Measuring and improving recall in efficient language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.896786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.079998Z digest=sha256:41d5b1059c59ce8d99bac78297ee5588dfc2bd058ccc5e1cf5ef47a699c24783

Observation af1324c6-0fab-4832-a706-76dc132db978 · outbound

This paper cites Fast attention requires bounded entries.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fast attention requires bounded entries

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.880702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.085462Z digest=sha256:e89d30baddaea18168b200543f19fa2478e361a151f77bc51660f51dd00b130f

Observation 8a8047c2-678c-45a3-a1ac-422c5871ea0f · outbound

This paper cites Fundamental limitations on subquadratic alternatives to transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fundamental limitations on subquadratic alternatives to transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.865671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.090223Z digest=sha256:5344b82f16f88d3b1ce6c915c92e18b54ec270e9ed4c52d286eaa0ebf7f04508

Observation 7945a37a-e689-418e-b8c0-acba451ad33f · outbound

This paper cites On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.850871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.095765Z digest=sha256:c96bf357a707f5f63ff0808d94c21b3f75ad9ab22b4b4c12b93f2d98f60cb19d

Observation fa65ef42-f46d-40c2-93c3-37b94a062016 · outbound

This paper cites Separations in the representational capabilities of transformers and recurrent architectures.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Separations in the representational capabilities of transformers and recurrent architectures

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.835057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.100144Z digest=sha256:2f6d45965530806ad2d0c99ba01a29aeca39e6aba98bcacd46056bfb1d9c6e49

Observation 024fca3b-7758-4ed9-9ec7-c685edee9285 · outbound

This paper cites Longformer: The Long-Document Transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.104615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.104615Z digest=sha256:f2470c4d66a96b9cc84f26c2cad86ff1380751e748baa1b645d3c3982d69a21e

Observation addbc6b7-327f-4df2-a6d5-7a5383a7a3a4 · outbound

This paper cites An exploration of hierarchical attention transformers for efficient long document classification, 2022.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An exploration of hierarchical attention transformers for efficient long document classification, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.819050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.110449Z digest=sha256:015d2c3d210c7875a8a40ca1192ca30a5a964603e94697410f737b89bcc41285

Observation d63ca6e3-07d2-4db9-ada5-861a5bfe1370 · outbound

This paper cites Colwell, and Adrian Weller.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Colwell, and Adrian Weller

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.114975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.114975Z digest=sha256:5ca8f4b1379e00ff36fb41743ec0d695b2fdda368fda17312f850140b4e64134

Observation 208bc6d9-2b45-4ef8-adae-3785c449204f · outbound

This paper cites End-to-end object detection with transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones End-to-end object detection with transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.794430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.119593Z digest=sha256:10cf5eac3f5bd88adc79e55c9262d64a3487f5d6cbb98a71379f1f8e1700f216

Observation c0d9153c-e51b-4936-abef-f0a471cc6c8c · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.123953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.123953Z digest=sha256:e9b0cd6b281a29a8ca01df1f451a16dd7bc8c21d96dcc9e6f231e4655df195ff

Observation 74419346-96a3-4ec1-940e-a64dd8201936 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Fu, Stefano Ermon, Atri Rudra, and Christopher R\' e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.779890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.128702Z digest=sha256:f09575466dff7ec33ac546648bc07d907a8f4394c00bc810913427edaa9e2225

Observation 02e50453-d976-4a8e-a032-008066dfdcaf · outbound

This paper cites Etched is Making the Biggest Bet in AI.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Etched is Making the Biggest Bet in AI

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.765080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.133124Z digest=sha256:bc01ac562d73098dd973a917e5611f71a2f58129a5a507c4302abd1bb860b646

Observation f4cb5a0f-b75d-436c-996d-6573336e49c3 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.138043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.138043Z digest=sha256:cf4c796bdd791436f2784035288f208e4730d65b39e89844fd47a483dfb2bb18

Observation a02b2ac7-3259-4954-8b3e-8da43426ddb9 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Theoretical limitations of self-attention in neural sequence models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.748865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.143413Z digest=sha256:c0493497914eea022a8159366021dd2f31d006797c47ab88f2ba723d7d84cddb

Observation 38ea6915-586e-4ae5-941a-ada67e86cc74 · outbound

This paper cites Hyperattention: Long-context attention in near-linear time.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hyperattention: Long-context attention in near-linear time

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.734346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.148089Z digest=sha256:ea8e81141df00d55f40a6a2aa385cc8fa3b33450be085ef1f6e0eff83dee82df

Observation 08d07ec7-a333-4c40-8c90-2aa45499d470 · outbound

This paper cites Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Computational limits of low-rank adaptation (lo RA ) fine-tuning for transformer models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.719060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.152253Z digest=sha256:351e1fe559bd6e14b00fbeb2c4bfeb429ecee438762fa64efac356ce4dc68217

Observation fe1866d1-3cbd-4e71-9b22-14369cd71c7c · outbound

This paper cites Multilayer feedforward networks are universal approximators.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Multilayer feedforward networks are universal approximators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.703811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.156366Z digest=sha256:65f32b482021b81c6f12d655311fabfdd2b4d9128c42e153ec4cc1664c491248

Observation 38e00788-d632-489c-995f-418982c0cfe1 · outbound

This paper cites On statistical rates and provably efficient criteria of latent diffusion transformers (dits).

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On statistical rates and provably efficient criteria of latent diffusion transformers (dits)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.689004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.160515Z digest=sha256:7f402babfbd6b27b5e64aae320a328b2b334c7a5e2f235e3c6ad7d9b3648f2e7

Observation 7cdac334-4f01-435b-8d57-51f8908b4925 · outbound

This paper cites Kakade, and Eran Malach.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Kakade, and Eran Malach

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T01:12:48.165303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:12:48.165303Z digest=sha256:1bc09262ab2373adf7806a92ac472c94beb6491f5245b89ae22b41cc600fe109

Observation acebf52c-beb3-4a63-ab6c-4ac01dd9cf27 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones An image is worth 16x16 words: Transformers for image recognition at scale

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.663817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.170867Z digest=sha256:d962e860e924b601352d5d4aec9acbb953d4c5674c0700b87b8044ca91059309

Observation e9d094b8-0055-414d-9fb4-a96198d7feb0 · outbound

This paper cites RNGD preview: The world's most efficient AI chip for LLM inference.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNGD preview: The world's most efficient AI chip for LLM inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.650041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.175070Z digest=sha256:3349a6c084de176e70d3fd817b99f2dceb97415e18dc839cd10ec70f6fed8b5d

Observation d80f281d-47c8-4a70-a344-5a40cd0df65d · outbound

This paper cites Reformer: The efficient transformer.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reformer: The efficient transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.635829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.178891Z digest=sha256:f66612c61970774a5e71cbf63eba101e00c0d178ff742e1c02eab252bdcb52b4

Observation 278afdd6-e958-4ee1-906d-af3f472e816a · outbound

This paper cites Polysketchformer: fast transformers via sketching polynomial kernels.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Polysketchformer: fast transformers via sketching polynomial kernels

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.622136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.182787Z digest=sha256:221a30e29e1201875cad20255feec14af9400d4eeb99d559d71c3c765b497017

Observation 48ed1d42-b864-410a-bf0b-22644b5e50be · outbound

This paper cites Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.608236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.186811Z digest=sha256:d8cf1505e6b09e60cb192c73f1de873067aec61a336eb9bfe921915f117e8aa5

Observation ed2960ec-a1f8-4d69-88f3-f2648a08b744 · outbound

This paper cites On the expressive flexibility of self-attention matrices.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones On the expressive flexibility of self-attention matrices

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.593022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.190765Z digest=sha256:f1bdd6724477d08e9613e65678fa65c73588c136b12cc2a5473d84e33a796802

Observation c29f4383-d564-47e4-b9a4-2b207dacee37 · outbound

This paper cites Hierarchical transformers for multi-document summarization.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for multi-document summarization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.578990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.194683Z digest=sha256:07fcb49a71fb2dde9dacb125ebbf1582ffc849b0ea989698e1891be8d8364eab

Observation d90d4e61-9a47-4870-896a-c408cef07b8b · outbound

This paper cites The parallelism tradeoff: Limitations of log-precision transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones The parallelism tradeoff: Limitations of log-precision transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.564989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.198545Z digest=sha256:c97e224f7ba146bc1751765676ec586cb6efa538f76fe5686982af5df656e061

Observation fc1f91c6-6d9a-44ea-9d52-57cf35de5df0 · outbound

This paper cites an unresolved cited work.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:12:48.550241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.202818Z digest=sha256:a6d9a4545c4f1893167db8736a21deec4298dc0ca3f93e028a35cf087523338d

Observation 3bb1e02f-d1fa-4f28-9bbd-3aa12a639221 · outbound

This paper cites Language models are few-shot learners.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Language models are few-shot learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.535978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.206861Z digest=sha256:301df0898c1baac3a3411d47efb199fd9f8d3163545361afd03d0b38d6bdeaee

Observation 3b258ad8-d2ad-446f-badb-12896773279a · outbound

This paper cites Hierarchical transformers for long document classification.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Hierarchical transformers for long document classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.521713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.210920Z digest=sha256:ef43d79974d0ffb4b323821ce594bb60204679d89231216843da896b24e433d1

Observation b661e8a2-a0a8-4843-add4-4105573376c4 · outbound

This paper cites Representational strengths and limitations of transformers.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Representational strengths and limitations of transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.507820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.214838Z digest=sha256:71669583a9906f1867e5564819a1f37cc5b9cd2546aeaefdcc6be8bd677c5875

Observation 3ef4a2f6-b742-4a37-a5c6-ff8c5eba978d · outbound

This paper cites Transformers, parallel computation, and logarithmic depth.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Transformers, parallel computation, and logarithmic depth

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.493360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.218877Z digest=sha256:4f92e3cb335d48f84e7cf052437f2e0e612bd730de90f255e1a671e1993123ad

Observation 44e1558a-3115-4c57-9557-a69b1b43a757 · outbound

This paper cites What formal languages can transformers express? a survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones What formal languages can transformers express? a survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.477544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.223004Z digest=sha256:653ead64ea190d4ed4d9ccbabd74818070a1b875a2a56358216caaaaef27a46e

Observation c1d55e78-58a6-4b59-9ea1-a56d3a76e0b8 · outbound

This paper cites Efficient transformers: A survey.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient transformers: A survey

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.462466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.227406Z digest=sha256:e3d11de1c9271e497b545ce1b73dfa4c51a24a7433f96f6b8a240e6c4804a23d

Observation 9445936f-2977-44c6-af2d-d60cdfd6cd04 · outbound

This paper cites Schmidt, and Stephan Peitz.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Schmidt, and Stephan Peitz

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.446971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.231424Z digest=sha256:a79c99ed4fa2825f30a33027d3d4105b31a8b7c23419ea07ccdae78901143504

Observation b3536e4d-94ec-4bcd-bfec-8e67154f4f8e · outbound

This paper cites Gomez, ukasz Kaiser, and Illia Polosukhin.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Gomez, ukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.429843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.235774Z digest=sha256:9644e9c285abb4a56dd23925ca493fc6d9abcce71aa00f66a098140b8c74e009

Observation 80d2ebdf-30d9-43e7-b2e9-9935b9e57e94 · outbound

This paper cites RNN s are not transformers (yet): The key bottleneck on in-context retrieval.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones RNN s are not transformers (yet): The key bottleneck on in-context retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.414278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.240370Z digest=sha256:80714badbff9484b5afd954ee5dc4a97c4e5992c5cae69fb11c312e42db99f19

Observation de93e740-306e-4720-b64c-b46d5a14bddd · outbound

This paper cites LightSeq2: Accelerated Training for Transformer-based Models on GPUs.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones LightSeq2: Accelerated Training for Transformer-based Models on GPUs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T01:12:48.303846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.245834Z digest=sha256:b000d9bae0094494bfcc0862527700bda1b9b100134913f7744d21bdf0de4b91

Observation 7d215dfb-3d60-43fe-8962-399d38217616 · outbound

This paper cites Efficient streaming language models with attention sinks.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Efficient streaming language models with attention sinks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.398365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.251148Z digest=sha256:f156d033262f05673ed9487c682359742d1fc0945671e4a86a5a50c82c9d3b2e

Observation 12d921ac-cc04-4481-b458-297b95299f53 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations , 2020

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.382813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.255183Z digest=sha256:0865dfc57aec3d718ae6740660f2efad6ba0f65e21046a0f72dc1c5017f226cd

Observation c1d80b37-089f-44c1-868f-5cd1ac530dc0 · outbound

This paper cites Reddi, and Sanjiv Kumar.

Two Heads Are Better than One: Simulating Large Transformers with Small Ones Reddi, and Sanjiv Kumar

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:12:48.366538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T01:12:48.259338Z digest=sha256:64f7ef6d425f52e14bd4c777d4c7fa5fefcc85651650873dcc8100949c0b3b45

Pith citing papers

Observation 795242bf-3d0f-49a8-96ad-9f827583f903 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Two Heads Are Better than One: Simulating Large Transformers with Small Ones

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.232597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:f37f9f608728a1d4f45bdf33fe6a51c63ef0b2b09c7ba3f1fd9455d06fe5e1f8