Pith. sign in

Paper Citation Record · LEDGER

Primer: Searching for Efficient Transformers for Language Modeling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2109.08668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.08668 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:04:38.115083Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:01.480491Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e1eadf9-6752-4c38-8a22-e808dfd44f84 · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models Primer: Searching for Efficient Transformers for Language Modeling

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:14:25.861443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:6b594b6fbe499747b8302cc781fe24b753ef95e18efe20aafd18e762f2db2138

Observation 152da4b4-8478-4165-8a0c-153ba16efdd4 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Primer: Searching for Efficient Transformers for Language Modeling

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.140017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:687d3de78846bc97510ce93e9c7c6f73229128f01ca632d4f6eb00713b3f0139

Observation 7e631a03-93fe-489a-ae0b-77f276dc3fe6 · inbound

Fast Inference from Transformers via Speculative Decoding cites this paper.

Fast Inference from Transformers via Speculative Decoding Primer: Searching for Efficient Transformers for Language Modeling

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:52:00.193939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T22:52:00.101612Z digest=sha256:7b83bf3ab5b129ec42232ad3464835ec6ed72dba82cd30a6fc5e2092ee6addef

Observation ba483445-22e1-47fa-bc2e-426623c4cc27 · inbound

TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations cites this paper.

TriADA: Massively Parallel Trilinear Matrix-by-Tensor Multiply-Add Algorithm and Device Architecture for the Acceleration of 3D Discrete Transformations Primer: Searching for Efficient Transformers for Language Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:38.115083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:38.115083Z digest=sha256:affa918817ac84f950b9ba0f607f7b531ac8d09bffc805a628f0cee33052a845

Observation 9b6b6d0c-110e-4537-ad71-4e3d70ec4747 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Primer: Searching for Efficient Transformers for Language Modeling

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:19.008474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:19:25.483640Z digest=sha256:53d57edd99384edbe2e529a7ba20da5c82ece24ec67f1c2f9ad3d1a44b37e785

Observation a97283ea-c4a4-401d-a35e-43b2ee9dd29d · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity Primer: Searching for Efficient Transformers for Language Modeling

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:54:51.152308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:54:29.436149Z digest=sha256:3a34f5a9b643aafdba0bac0421342fef383baf01150fbdfcf0df18172fd43417

Observation 2dc3b978-431c-4f0a-9962-9642290226d5 · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Primer: Searching for Efficient Transformers for Language Modeling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:52.961943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:52.961943Z digest=sha256:1b218689d50c7fc4d141500f1b8f451eb0caf89aa4780fb124cf12858c2e1838

Observation d9ce9369-a7ab-4048-ab75-8ec05532f67d · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence Primer: Searching for Efficient Transformers for Language Modeling

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:40:42.503282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:58d1e3a58ad106c37b5da15105e7d5d6562217c66ee68f397e88bd9dfc8e97bf

Observation 31333250-5087-409d-9d92-d90c28a5c707 · inbound

From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums cites this paper.

From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums Primer: Searching for Efficient Transformers for Language Modeling

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:47:33.252670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:42:48.141279Z digest=sha256:efb5e778bbd551544372c99454248c08113d4c1db6edb188b3e05bfa952e4fc6

Observation a92a4ad1-0ee7-4fac-9646-899228e1fa5e · inbound

Three-Phase Transformer cites this paper.

Three-Phase Transformer Primer: Searching for Efficient Transformers for Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:24.629483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:04:23.797546Z digest=sha256:83fcf575a2eef8fb5ece8235b369ee8a119bbf29399c9046010d3be69a13919a

Observation 81a24bda-88dc-4f43-8406-f21e617dc975 · inbound

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity cites this paper.

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity Primer: Searching for Efficient Transformers for Language Modeling

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:26:12.854349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:07:18.278784Z digest=sha256:a290e1d9657ef9a4b86061e2f51c907770fea304040c6a2884261d42433bc826

Observation 9af5fca8-be7c-47d2-ab7e-c05513672db6 · inbound

On the global convergence of gradient descent for wide shallow models with bounded nonlinearities cites this paper.

On the global convergence of gradient descent for wide shallow models with bounded nonlinearities Primer: Searching for Efficient Transformers for Language Modeling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:31.411140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:51:08.871267Z digest=sha256:7eae9be6ba58e487cd59c566a8528527700c4910363c8117ba5009d8369748d8

Observation 9e074bf5-44df-4e0c-8516-be3272f0c9cb · inbound

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes cites this paper.

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes Primer: Searching for Efficient Transformers for Language Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:48:19.650855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T13:46:32.405079Z digest=sha256:48451d95b479e08f7cfa1420843849d7a86e3c0ce174564cf2a99823e10aabc1

Observation 41f62084-2ea2-415e-bf29-d466206687d0 · inbound

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes cites this paper.

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes Primer: Searching for Efficient Transformers for Language Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:16:19.681494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-22T09:15:32.395442Z digest=sha256:df0637588e4200f21d566a4d7089287b15c6d6c96c0d3a912b6df0d511d3af9b

Observation 03ec595f-8600-461a-b6b7-3a795559c30a · inbound

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training cites this paper.

Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training Primer: Searching for Efficient Transformers for Language Modeling

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:14:01.482003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T23:10:47.199537Z digest=sha256:954fb035f2f0cb9fd1122c7491eb116ad3a33c5394ef7da94f520b97efa17c04

Observation 430ad754-b9ad-4167-82e4-4d64d919e1c6 · inbound

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks cites this paper.

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks Primer: Searching for Efficient Transformers for Language Modeling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T17:31:49.346142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:31:49.346142Z digest=sha256:bc9086da39234364cf87939543e117f879133e2ced8c1a7c595dcc439289f94b

Observation 9f1d067a-18da-46a7-b512-07118b323b7f · inbound

Domyn-Small: A European 10B Reasoning Language Model cites this paper.

Domyn-Small: A European 10B Reasoning Language Model Primer: Searching for Efficient Transformers for Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T14:07:52.153859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:07:52.153859Z digest=sha256:35a8cdaa27e45cb4b0b301027186bea8121a513db5af8c1d7671fd2b3bc95129