Pith. sign in

Paper Citation Record · LEDGER

Full Stack Optimization of Transformer Inference: a Survey

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2302.14017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.14017 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:08:59.936655Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:38:49.031648Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 71340de5-fc2d-4173-b6c7-cb3a88a7f950 · inbound

Foundations of Large Language Models cites this paper.

Foundations of Large Language Models Full Stack Optimization of Transformer Inference: a Survey

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:58.925228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:14:58.925228Z digest=sha256:311a8069ff5c88b49f35d0e3cc441d2dc14a9bafa27d473f35b1108a49883c4d

Observation d2b4ef67-ac41-40cd-8104-def8e9c42dd5 · inbound

Pushing the Limits of BFP on Narrow Precision LLM Inference cites this paper.

Pushing the Limits of BFP on Narrow Precision LLM Inference Full Stack Optimization of Transformer Inference: a Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:22:18.802764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:22:18.802764Z digest=sha256:279673fe7222ef97505087f974fd17d62d192cfb0cfd97384556975a2b88d3e7

Observation 1dce0f8e-e1f0-4d76-84ba-eaf82b5e4aa3 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Full Stack Optimization of Transformer Inference: a Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.547913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.547913Z digest=sha256:fb4345219b1e1bfda18c9de3f9dbc741e62c3100719aa2412db6ee0fa5c7bdbe

Observation 0c6e26bd-44e7-482e-a6d7-a681543d02d7 · inbound

EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge cites this paper.

EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge Full Stack Optimization of Transformer Inference: a Survey

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:59.936655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:59.936655Z digest=sha256:280907bf64012f0ca285852ec1ad7d44c68f1f636eb6b03f937697a65c6ec7eb

Observation 295d31ee-b35e-4921-8889-44aef5a7985c · inbound

A distillation-teleportation protocol for fault-tolerant QRAM cites this paper.

A distillation-teleportation protocol for fault-tolerant QRAM Full Stack Optimization of Transformer Inference: a Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:55.019799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:55.019799Z digest=sha256:12b19838fda4a53ef752f8691153d87cdc747ad8fd3352b2c601a780fab0f37e

Observation ff7dada2-ed6b-4b18-94ab-09072f391012 · inbound

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations cites this paper.

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations Full Stack Optimization of Transformer Inference: a Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:28.605678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:28.605678Z digest=sha256:4023dad47387d48009327a17e262c6c0b9fecd0b81afd6f7ae010928697af8f5

Observation 250be190-98cb-4066-b14f-2ead6dde0bdb · inbound

COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives cites this paper.

COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives Full Stack Optimization of Transformer Inference: a Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:29:57.252635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:29:57.252635Z digest=sha256:3ccf4c91c7f421e8bcebe44154f366b3b35e2a91cc30fe84a242882bca6c492d

Observation 899656f8-c19e-41f3-aed5-d9ee1e929726 · inbound

vAttention: Verified Sparse Attention cites this paper.

vAttention: Verified Sparse Attention Full Stack Optimization of Transformer Inference: a Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:21:07.403398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:21:07.403398Z digest=sha256:d1d2260b33cb7482e59b20d5edefb9184268c368057243e22b4ca2be7153802b

Observation de6ce19a-7c98-411b-843c-8cf7c0e94d6d · inbound

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs cites this paper.

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs Full Stack Optimization of Transformer Inference: a Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.960085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:33:14.766545Z digest=sha256:b777593c548ddcbf228c2a3b5b3e0fe7e80d724fac43fba4de7cf2ddf870adb2

Observation 8cb40e57-c913-4f11-b89c-d47ac3da4d64 · inbound

Learning to Remember, Learn, and Forget in Attention-Based Models cites this paper.

Learning to Remember, Learn, and Forget in Attention-Based Models Full Stack Optimization of Transformer Inference: a Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:06.492108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:06.492108Z digest=sha256:167cea979140a2b0cee15103bf76236ceda72c21f6c443c59d6e67c7ff4c3794

Observation 6c75d491-37a0-43e5-a01f-73441826f58a · inbound

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures cites this paper.

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures Full Stack Optimization of Transformer Inference: a Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.486747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:33:59.777818Z digest=sha256:ab357b0f6fba618c4916ada13a427ce845bc43cd87989e1290a619f36e60f054

Observation af522d9d-bae7-4c65-94db-fd8fefdb935a · inbound

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation cites this paper.

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation Full Stack Optimization of Transformer Inference: a Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T23:32:59.467653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:32:59.467653Z digest=sha256:e6110ea4215563a9b42d657a6371439d7ae54915ccf6d9fabc018c597d04453c

Observation 26e3db36-ea6b-4493-bfe9-30bd271a7079 · inbound

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models cites this paper.

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models Full Stack Optimization of Transformer Inference: a Survey

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:16:01.588643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:04:49.793158Z digest=sha256:de0b20bf886a5f86b15806fe7c8fae0a4a2e479e08b3844ef1b21b6f05a58e05

Observation 431c1525-8904-44cc-9e97-8d137d5a0425 · inbound

CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration cites this paper.

CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration Full Stack Optimization of Transformer Inference: a Survey

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.117402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T07:56:53.425178Z digest=sha256:8602e35bd5d11b14573dd2725cd28d51bbec4717263c9e142e4635b689804812

Observation 9e9887d7-95d9-4a33-a239-09e47bd96931 · inbound

Edge-Inference Governors Need Memory-Clock State cites this paper.

Edge-Inference Governors Need Memory-Clock State Full Stack Optimization of Transformer Inference: a Survey

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:38:49.033705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T02:44:49.539266Z digest=sha256:1ffb22ee724eb5670725b329f1b42011273eda2d90591e0611574f724d833cc9

Observation 5fd2401b-6fbc-46ee-b85b-96ae1df2ebdc · inbound

Edge-Inference Governors Need Memory-Clock State cites this paper.

Edge-Inference Governors Need Memory-Clock State Full Stack Optimization of Transformer Inference: a Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T13:55:19.148570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:55:19.148570Z digest=sha256:b4fbf6325b7c18dd555b97d0478150ebedfa8b4bfa87860a3257c438b061e882