Pith. sign in

Paper Citation Record · LEDGER

ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2303.08302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.08302 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:57:41.320727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T00:56:24.897034Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f26d289c-09a8-4f5c-a6f8-ce0ad0956639 · inbound

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance cites this paper.

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:45.442680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:56:45.348973Z digest=sha256:0e6d7f12d99850cf0d9a7cfc21a5bf9e47a41d287012743df61b25b2697ace89

Observation f06b8139-c80f-4fba-aeb4-50b171d3cb12 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.241090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:99dc44119c94d77ee1da3056307859cf84723a8373ef08b2c8fd12912e3d7698

Observation 5ed274c0-c79e-4d6a-8a69-55a310c6f575 · inbound

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation cites this paper.

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:41.320727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:41.320727Z digest=sha256:16db0910a2988a6597d2c3d1b92ef7789e09791d5ba00c73cac1b718a206e3fb

Observation 25797477-e26a-421e-9ef0-e44c03e8bedc · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.467026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.467026Z digest=sha256:3123a30d44a28dcf068f658830ff0fac4d91a0c14a83b4066fe147c927b1ffb4

Observation d49f8fb0-a2f7-488c-890d-ab937df70595 · inbound

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks cites this paper.

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:13.303967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:11:13.303967Z digest=sha256:1559a6b615ae9ef00ae49eae59c65e63098ba79a13ecd14889e332acb1b210b3

Observation ef2bf2c0-6452-4e33-bb6e-4c5c73e30e2f · inbound

Enhancing Model Privacy in Federated Learning with Random Masking and Quantization cites this paper.

Enhancing Model Privacy in Federated Learning with Random Masking and Quantization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:10:53.398370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:10:53.398370Z digest=sha256:7ad7a3383fae754ffed86663b7c2e47fe90ce66f6aefffd32b436594b31e779e

Observation 8566c60f-50ec-4410-94e6-6776a6cd901a · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:50:30.273755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T18:46:04.926179Z digest=sha256:5cccbf7b5835a20851a4f9c1661d218f2c141320bb3ea31b14376afbc4e12d7a

Observation 9777d95e-76ab-4c39-b98f-ae42929a6f48 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T23:21:42.110167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:21:42.110167Z digest=sha256:4df6707dfdb1d6a7b93b32448059a8f6940af97a9b2f62ae9846d9ac4b951907

Observation 310bac23-4950-49a3-a4ed-0b616205c63b · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.415115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:deaa924e6aa4517464815dbee6fd08bce9ce6a93db897411de0739ae248334ea

Observation f05640c3-6400-4fc2-b46b-aacd0f9b81f3 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:36:19.950190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:36:12.915133Z digest=sha256:e6b20d3e33acf6931a1b99df2e128d16b5509f1811f979038e5dd5189a49b254

Observation 03866696-993d-496c-93ff-04e990ad6929 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:30.196922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:29:14.545746Z digest=sha256:4a2d2d963f04c5f087ae45f8eb214f0d7858bc06796a1ebc82f6d4cfaed386ab

Observation 4246816d-2079-4874-b669-88e413c3a234 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:50.166136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T07:57:49.746594Z digest=sha256:b63c92e157bdb0cd7f8179aeaada63f6e8d68b7f477b6c49b617c2d861c388c2

Observation bc7bccb3-9cea-42ea-8cce-3c2bf492175e · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.633191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:9f42893ab19a598b717d735ee38e78b6db027fe4e3ac2ce10f77646f7fa3b42f

Observation 9e9c1558-3a9b-4696-8b17-2112027836db · inbound

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression cites this paper.

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:12:34.512202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T19:10:58.845006Z digest=sha256:72d93a322f6da4d8d88c3920c280b87b47b6d5f73d6752d0b829407c6ba65a0e

Observation 07652548-2994-4d7f-9f53-23a64a978848 · inbound

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation cites this paper.

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.396687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:28:14.160341Z digest=sha256:7f771b84cc435d8f68913b62beaf8a712648e1bad1810a4ace8a0c4f0924e8c7

Observation c504d84a-0b2d-4614-a750-898e16cfb00e · inbound

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization cites this paper.

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.899516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T13:07:53.001390Z digest=sha256:f0959e9b4377615ff8c42e4fb9e1294fb93a0e16cc73d14ef4d86cf73f634237

Observation 66defc3c-1c96-409c-be9e-3130e323c914 · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.125011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.125011Z digest=sha256:c1600a5f74ad99822a6843882879a15558022fb2a70cfc4b96e710cb54468821

Observation 5b0249f8-bee6-47b8-9021-6f42da1652e7 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.948303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.948303Z digest=sha256:343d5317a26ab2e285fbfb3c0d1aa86907f212b2ae78c82aebd9e8bedbdb0b99