Pith. sign in

Paper Citation Record · LEDGER

ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2206.01861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.01861 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:38:22.309181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:16:16.046751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5d495910-c8b9-4bfa-b361-25c92364875a · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.185243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:b5a54ba6b1ef4dacc83eb520b05c949096f4d08ee9d4b67dc9c5bff6cf24935d

Observation 03b4f879-9892-4bcb-abf0-9017dbf482d2 · inbound

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers cites this paper.

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:18:35.274980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:18:35.153078Z digest=sha256:57a5796b03ef27ee8ad64b813f492208a6be59a711fc1db1f8f51d426efbe770

Observation 3b7bbd66-67a3-49f0-b4ba-f25606e2f5ac · inbound

Accelerating Large Language Model Decoding with Speculative Sampling cites this paper.

Accelerating Large Language Model Decoding with Speculative Sampling ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:29:36.307343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T07:29:36.205026Z digest=sha256:774e58d94c9f7804e85342c4421619529eec417f5dde91adfd5d56f8ae8b4b8e

Observation f6589bb8-472a-4d55-b1fb-2a71b309c258 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:29:53.790343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:fc9f3baa35f012de49e62febc1d6d47beb4528e0e0eb7814390c9f2dfd7cb4e4

Observation 419f3a01-d77c-4b52-9378-02df1fff1d21 · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:29:11.224966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:54c0f067e68b00e3d4a3405c488bc9b67d4d362cce5c0a3707fc11dc8f7d627f

Observation 29d83d7a-d3a3-4d48-b1df-85a29bb200fd · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.460651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a8467ffc89278b480614aa784da3a892dc2d9b8887dad8af7bd647dbc7916335

Observation 405b02e5-3090-4efd-b4dc-2c5bd61b737b · inbound

BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration cites this paper.

BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:18:05.575588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:18:05.575588Z digest=sha256:3e864e83216a4ac116f4a4685a5baeb397add112e8da3650b99a750398b88783

Observation 9015d2e5-5189-4571-abb4-2871b0e365e5 · inbound

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem cites this paper.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.258416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.258416Z digest=sha256:96a03729e57ed0f768a0c8c5e9e4e234bf286395503805c9442cfd6e6bb9afd6

Observation 4cc0df57-4fc6-4af5-9030-8a7ae42ce0c6 · inbound

4bit-Quantization in Vector-Embedding for RAG cites this paper.

4bit-Quantization in Vector-Embedding for RAG ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:12:58.837909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:12:58.837909Z digest=sha256:cbe9eb33b9487fe7e4058f45f73c0174d095f825ad68746f904e21fe4bf2a8ff

Observation d1c187c5-2f9a-4c8e-b164-77642734c8de · inbound

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting cites this paper.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.988135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.988135Z digest=sha256:6a7c42ed072e0793c2baae3920aac52ee30c1adce157505e4af0c1cb65b17ecd

Observation 6a49b61a-c3c2-46dd-a9a7-c8c2c44b479b · inbound

Compute-Optimal LLMs Provably Generalize Better With Scale cites this paper.

Compute-Optimal LLMs Provably Generalize Better With Scale ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:22.309181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:22.309181Z digest=sha256:699556c1620408a6d8c8d4b149adbd4875a1cdda4ec3aea0eea5ace6b20de50f

Observation 45ec0f7b-73cb-4e47-bc37-5d2b3f5d44d3 · inbound

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference cites this paper.

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:07.877611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:07.877611Z digest=sha256:0b814db0b1bedcadd8c83d4bbc94d8fd4ee5dc93aa8ed93f737d24303404df76

Observation a138f024-ee03-4bc0-97c7-7b1d9e2e4d83 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.045802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.045802Z digest=sha256:d2c2016b4b751b46c542b489df008557b5c82f2eff1da666fa2b54b9b6ef64e7

Observation 8be7de8b-4369-49d6-98ba-720147f5f4b1 · inbound

ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models cites this paper.

ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:06:29.118203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:06:29.118203Z digest=sha256:739ccc70520f460133b03fb63842cc4e4fe0f4ae1c727c9369a6c138b6d2fa81

Observation 934d023b-00d6-45e5-a64e-f1d9d38e76e3 · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:47.848505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:47.848505Z digest=sha256:1115486bc828df9c7aa9d866d197bc65cd36651fe26513c613844911a7c2e73c

Observation a3f54d90-0c07-462f-b783-f990656e6967 · inbound

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models cites this paper.

Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:31:08.335975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T17:30:43.808662Z digest=sha256:46aedf23256e0d21adc5fcea7d956478689fc7a4a6b1217bb977125ce44316c4

Observation 38c47fdd-5b79-4043-b192-f7df2b5a3817 · inbound

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models cites this paper.

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:25:29.242563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:24:53.745784Z digest=sha256:d301080d5faffcd38224a45b5c9b5e5928f71ebd2b564a921aae6e32e9b694ed

Observation b3d0d9e1-e4d3-4c83-882e-1e7635f2e040 · inbound

Motion-Compensated Weight Compression cites this paper.

Motion-Compensated Weight Compression ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.440211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T12:58:27.637822Z digest=sha256:5c0dbdc7cd5b89b86e601a566bac39ef31e2225cc8c2592caa66e3b54cd8add7

Observation 2687967a-64b3-4698-8ad1-f18a8546b249 · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:16:16.048491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:c9f9d4fd11f13c30ba365da0367ab666f170ed511415ff35c3ac10b409bf67b8

Observation 6276d88e-1bee-4694-b241-3aaf305918bf · inbound

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights cites this paper.

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:51.939462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:51.939462Z digest=sha256:6a6b7086a79000ca2fd2b5c46c39c9fe80cf1efc5736c60e6881160024257336

Observation d81b60d8-44d9-4fd9-a26c-ae746f334db4 · inbound

Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra cites this paper.

Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:55.878782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:55.878782Z digest=sha256:ffcba7964e65388f702c57915da0d5577652198c5350f742adfaed65e432dca9