Pith. sign in

Paper Citation Record · LEDGER

Understanding and Overcoming the Challenges of Efficient Transformer Quantization

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2109.12948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.12948 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:05:12.843611Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:36:56.029956Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45a3d574-8780-4186-9f90-ab8356ad0364 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.041677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:e96c73bb42995df6f3785fa4996938fb94810e87fd9dde8e3db0e407332a1e02

Observation e82989ce-b32f-437b-a7e9-e613644e434e · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.968703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:d38d0ece3834ebafd7ad1e8a3726e6375e64e77b801512f76c815088270bb920

Observation c40f9115-53c2-429d-8db5-acb26df7374e · inbound

ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits cites this paper.

ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:08:28.638542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:08:28.638542Z digest=sha256:8feca2b7c4365cda5dc938fc5777f4769d163414865dfcf470bcce3860016f60

Observation af142f82-111c-4f84-b458-5ef43c77c9a3 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.546719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.546719Z digest=sha256:41004d308db70c768bc295d03d69c00caa35c0b3cab1aa08c1cb20d2dcf3d7b3

Observation de8badda-8ef8-4218-9f78-7ce7f126f326 · inbound

DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models cites this paper.

DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:07.011927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:07.011927Z digest=sha256:79fb481809ca7e779592b1e14d80550530f7bb83667fed25bdbac39dffab31d8

Observation 8f4fde32-7388-4520-90fe-067769996c82 · inbound

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models cites this paper.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.843611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.843611Z digest=sha256:5a1d90018f6e41fcfcc68092c163068db239f3456afa80b16a46029515083e0a

Observation a077caf8-59c0-48e3-b2b3-a40fc32dc6c2 · inbound

Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques cites this paper.

Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:02.061610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:02.061610Z digest=sha256:d1fc6f3c36679aa8dac8bef361bd754019ed8e181b8d8134e784009e91c52b6e

Observation b612e36d-a5c9-4395-b7d3-b36f369f8b7e · inbound

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference cites this paper.

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:07.616176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:07.616176Z digest=sha256:0be5a4d54b156da7aca42437ec360177b042d03cd0553fbd5bd1e7476768ef36

Observation 46096209-e9cd-482a-8745-83417ccf525d · inbound

Post-Training Quantization of Generative and Discriminative LSTM Text Classifiers: A Study of Calibration, Class Balance, and Robustness cites this paper.

Post-Training Quantization of Generative and Discriminative LSTM Text Classifiers: A Study of Calibration, Class Balance, and Robustness Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:14.109364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:14.109364Z digest=sha256:0e685c80ce0b5e51e31f7a9172dd023211c0265cfb29ed7ab413a0b34d7c0b43

Observation 49ff8c98-eafd-4f81-bb6b-602d93bc7237 · inbound

I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation cites this paper.

I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:09.346300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:57:09.346300Z digest=sha256:ae1b3adc8ea1267bc2e3d186d9a99b84b842386a1d4da7ee1978713e21f2996f

Observation 8653d2c7-9cb7-43bd-8f0d-526ea9e400b7 · inbound

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models cites this paper.

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:51:28.491836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:51:28.491836Z digest=sha256:3ee29d7d07169f1ebaf44bc928eb48516f5aa18c7bac614f88e50caa19aff02f

Observation 4d9a76ec-d298-49ca-97e6-b9bcb050be85 · inbound

MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition cites this paper.

MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:50.803740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:06:16.532479Z digest=sha256:ae41478cc9eb74ebd3a9148ad33a0e9a385c790f0a082673dee5af039609582f

Observation d3029044-f5a8-4af6-aa63-fc8efe78c83b · inbound

$\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space cites this paper.

$\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.031421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T12:35:58.613973Z digest=sha256:7e724b522b75e55c21016d026b786e10e82605c1b18f83e6b7a2425dd1c39e57

Observation c1bb3eca-1d65-4534-ab2a-c7672dd1587e · inbound

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference cites this paper.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.428735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.428735Z digest=sha256:69b86903732d63d34c20f309e16a0664d2f1ec688b693c1e3a977b5c81ec40c3