Pith. sign in

Paper Citation Record · LEDGER

Efficient Transformers: A Survey

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2009.06732.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2009.06732 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:48.847581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:26.546053Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e99b873a-8ee7-47ea-bd68-6329a042f71d · inbound

Deformable DETR: Deformable Transformers for End-to-End Object Detection cites this paper.

Deformable DETR: Deformable Transformers for End-to-End Object Detection Efficient Transformers: A Survey

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:47:17.015597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T09:47:16.915936Z digest=sha256:a36cc90de81e194dad4de8e5bbd64402cece5ce296ce00e008eccae6a535e11e

Observation 12edea71-84f4-4329-82d4-5de80f6a45f3 · inbound

Perceiver IO: A General Architecture for Structured Inputs & Outputs cites this paper.

Perceiver IO: A General Architecture for Structured Inputs & Outputs Efficient Transformers: A Survey

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:47:13.942394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T19:47:13.831505Z digest=sha256:65945001adc0c5c86bbb575d25e12fe88af4059d1247d1d6f1f6315cf8632260

Observation 7b2fb599-be62-4cc2-b64b-03d86bd32cef · inbound

Ligandformer: A Graph Neural Network for Predicting Compound Property with Robust Interpretation cites this paper.

Ligandformer: A Graph Neural Network for Predicting Compound Property with Robust Interpretation Efficient Transformers: A Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:24:27.496636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T12:22:23.989360Z digest=sha256:1ae21df7364abb27e4b93e557a52354a49c732c1154c8b164cf302654b4c1d98

Observation 5a3476db-bd8d-4fc1-9e49-5ef0d7a2f84f · inbound

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness cites this paper.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Efficient Transformers: A Survey

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:22:09.018695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T16:22:08.801066Z digest=sha256:6fa2728f99d4ae6bfc0ae64a5d454e4d2acd02f542e6119f92cf2f5e39936665

Observation 4b4a1bac-4f88-491d-9e00-a224064295ab · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficient Transformers: A Survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.482223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:ca3b4ab5c59c735141aab384192fc6597fb18554ce3b28214cf566cf452fb211

Observation 983eddb3-8be0-4bae-84c3-5ef51c7c26e1 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Efficient Transformers: A Survey

Reference 215

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:978ad73f30d07ed46c4b0b21300ba9fea714e1216c5a1bf176bba0cb0d79442d

Observation b2c62249-e745-459e-b996-bb62726b4fed · inbound

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models cites this paper.

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models Efficient Transformers: A Survey

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:02.299100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:18:02.254491Z digest=sha256:c436f22d9f0ee7e744e858ddc2d7b67d102f9e5f8a245ee50a22d9a1679e1860

Observation 9800e6cd-56dc-449e-983a-e102a4bd1f38 · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Efficient Transformers: A Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.448737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:f57c46fbd7f9201449a94a095d444b0f43f9db62ba8aea0df0b024dcc401c67c

Observation 5de5ed35-0373-454a-a6b3-dcf414336252 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Efficient Transformers: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.285388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:d2776cc7d41e768e09e0cef586a79862a805db1ee9e6dd81d532f452071cd6f9

Observation 4e49812c-ecd8-4cb3-817d-fde0fbffd9fd · inbound

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization cites this paper.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Efficient Transformers: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.847581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.847581Z digest=sha256:f39aabb90df63fa0cd24af08a728adbcd0876c87e32ca4c28ba2faec5b291632

Observation 63bf97f3-c7f4-429d-99be-880d433f5eb0 · inbound

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer cites this paper.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Efficient Transformers: A Survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.842653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.842653Z digest=sha256:4800a42c5b81c923c59da0f59efb8db9f1a83cf9b9ccf2c9138f5da4dfa6f341

Observation 10f83ea2-3167-47af-be9d-9e55322ee048 · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Efficient Transformers: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:31.651037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:31.651037Z digest=sha256:34625f768986920787f7f9e6485af1705812a34c503ad1f95fc84fd58dda28ba

Observation 7fffd0f1-c1c2-4559-aff7-3ba47c27c151 · inbound

Crisp Attention: Regularizing Transformers via Structured Sparsity cites this paper.

Crisp Attention: Regularizing Transformers via Structured Sparsity Efficient Transformers: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:05.668093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:03:05.668093Z digest=sha256:04bff38220c05e27e9777a5e517a535f9671578ee28b456d2975eab57a13b4a2

Observation a5368d7b-d973-4ea7-865b-76c75cc6c8f4 · inbound

Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling cites this paper.

Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling Efficient Transformers: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:30:01.075667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:30:01.075667Z digest=sha256:bd32672ec84cba03808872eb2260fac1a2b38b4a47263a151be79170ca4b6e01

Observation b87d149a-da18-4137-9bc9-9bb63f236dd8 · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs Efficient Transformers: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.024663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.024663Z digest=sha256:96d3015972ba15235a1caa81330c13dbafcaa18115b873b3c1d79a1f7f862c6e

Observation 65c88820-912f-4ddd-9b9e-4559d0d1bf43 · inbound

ARC-Encoder: learning compressed text representations for large language models cites this paper.

ARC-Encoder: learning compressed text representations for large language models Efficient Transformers: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:28:52.455570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:28:52.455570Z digest=sha256:4b5e9af7601c36d5a7acec218535f6c2ad1edf676bc4003cfd0e470b0f45643c

Observation b693897c-c815-4367-a8f5-c0746837edd4 · inbound

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch cites this paper.

TiledAttention: a CUDA Tile SDPA Kernel for PyTorch Efficient Transformers: A Survey

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:46:23.663710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:44:34.166709Z digest=sha256:cdff6bd13759a8d37ff81df689a2e3479d179b01d6270e39391bb2d6bdbd040b

Observation 35cf5da5-b212-4358-91a7-513006f24487 · inbound

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling cites this paper.

M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling Efficient Transformers: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:39:59.305385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T11:35:43.088803Z digest=sha256:ab33cc1a39eccd51c66b60a7f1ed14b3e5b8443cebf62737b7b30770e1412d05

Observation b2158cce-f87d-465e-86a6-3cd321ac5793 · inbound

An explicit operator explains end-to-end computation in the modern neural networks used for sequence and language modeling cites this paper.

An explicit operator explains end-to-end computation in the modern neural networks used for sequence and language modeling Efficient Transformers: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:14.897597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T22:42:10.197470Z digest=sha256:dc5cb19e18e3baa9dcef71dcd866121b7e60c0860de6b312e867b43af3f8ee64

Observation 0b00fe47-86b9-4927-9a66-7b1ac81dd775 · inbound

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents cites this paper.

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents Efficient Transformers: A Survey

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:49:53.685071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T08:45:56.550821Z digest=sha256:ea57e66f5b9e5a70b4963576652abe36f74e9ff9d42b22d14c05b8266cd3f9cf

Observation adaaeadb-cf62-472d-858f-a58caed80e87 · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics Efficient Transformers: A Survey

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.547593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:ff832d24f8db1803d0f9be3d5af8b7d8c10494e0779c6a43c9312793b18e78c5

Observation 73cedafe-d713-46d4-875d-b82042af100b · inbound

TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings cites this paper.

TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings Efficient Transformers: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:52:55.820093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:52:55.820093Z digest=sha256:434fbb59034e135f146ad5f9c1cd395d6d7433998f48a33ccc07a773577f6be1