Pith. sign in

Paper Citation Record · LEDGER

ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2303.08302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.08302 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:00:02.087433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T00:56:24.897034Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f26d289c-09a8-4f5c-a6f8-ce0ad0956639 · inbound

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance cites this paper.

FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:56:45.442680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T19:56:45.348973Z digest=sha256:739b08e53b0046960e3e1a9d1041f6b5700e05b3fcd9bb0f52130ecc296526d0

Observation f06b8139-c80f-4fba-aeb4-50b171d3cb12 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.241090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:deab437ca2838fc306ce434e7bd24127aadb48233c7452f9be8d1b4000795c6a

Observation 36f0a79b-4589-40f1-821d-06efef05c5b5 · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 212

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:46.750586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:46.750586Z digest=sha256:a6528a124a4b576ff73c9d530413e76f998e2ad23b11af162e3b83ca573c7b6e

Observation e13bac50-265a-49da-8a76-8133c812febf · inbound

MBQ: Modality-Balanced Quantization for Large Vision-Language Models cites this paper.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.305041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.305041Z digest=sha256:19d1d5669ed3ed934e7b8b3a1cc5504645e256dd9ab645d1d25e04172193458e

Observation fa207045-082f-4f89-a9bc-811e3d79261e · inbound

4bit-Quantization in Vector-Embedding for RAG cites this paper.

4bit-Quantization in Vector-Embedding for RAG ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:12:58.823451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:12:58.823451Z digest=sha256:49b270051ebb3ecfb065be9d40f106d59589c963fb75c7c663ed33971b52adb4

Observation 46c4b673-a8f0-45b8-96fe-44e33b05dfd8 · inbound

iServe: An Intent-based Serving System for LLMs cites this paper.

iServe: An Intent-based Serving System for LLMs ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:14.549116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:14.549116Z digest=sha256:3769ef4614b36c7f0992fd86f45921862017eb969bf7c7aa6745aee33351f672

Observation dae832e8-680d-4d9d-b77b-e0f4a1f3b475 · inbound

CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization cites this paper.

CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T23:25:54.734518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:25:54.734518Z digest=sha256:347e0b4eae30b067de4e90cf8ce0a67f724bc2a2d828a3a74217c71ec3973b7c

Observation f40fa136-bbbf-48b3-8d36-1fd779a4d639 · inbound

Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques cites this paper.

Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T01:00:02.087433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:00:02.087433Z digest=sha256:e9122ae067dfaf1f939fdddecf11e46d064ba9f927992b6186635e06c8e34502

Observation c69edea0-5bfa-4fec-bba0-34479dc12a9b · inbound

Low-bit Model Quantization for Deep Neural Networks: A Survey cites this paper.

Low-bit Model Quantization for Deep Neural Networks: A Survey ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T23:13:19.805481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:13:19.805481Z digest=sha256:384e7913175019ee2f75166bcb8cc0f59b6151a7ba1c739e69d7738ec23aef1d

Observation 95b31ac0-1b8e-471c-a3b7-6cd45ee75356 · inbound

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference cites this paper.

Resource-Efficient Language Models: Quantization for Fast and Accessible Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:07.881424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:07.881424Z digest=sha256:d65315d4d2c6a721066d772e6641830caa25c557ed3624e4c278bfbbac43ba14

Observation 5ed274c0-c79e-4d6a-8a69-55a310c6f575 · inbound

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation cites this paper.

Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:41.320727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:41.320727Z digest=sha256:7dd32eeed9a7bef0c1776483d3ece548845570cbd552f6e8eab76dac77ed72fd

Observation d7f95455-3a75-4a94-a33b-840c231f2507 · inbound

ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models cites this paper.

ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:06:29.123982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:06:29.123982Z digest=sha256:6a81e57beeae8fa664bc311f1d3806d4f91973f8a9f421e8b72fb1386fd90071

Observation 25797477-e26a-421e-9ef0-e44c03e8bedc · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.467026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.467026Z digest=sha256:54d435303a29938292ade4c7c50b5aafc75baa1fa2d953b54c36cf0925e314c6

Observation d49f8fb0-a2f7-488c-890d-ab937df70595 · inbound

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks cites this paper.

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:13.303967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:11:13.303967Z digest=sha256:82aaa7506d6ee466f48cdb41c2578429dac42fafb06751e7acdd0345e5237c74

Observation ef2bf2c0-6452-4e33-bb6e-4c5c73e30e2f · inbound

Enhancing Model Privacy in Federated Learning with Random Masking and Quantization cites this paper.

Enhancing Model Privacy in Federated Learning with Random Masking and Quantization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:10:53.398370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:10:53.398370Z digest=sha256:ff4e53be413e341ec09efebe52e42bf6feb25b94e4f1590fc70c929de8fd72b5

Observation 8566c60f-50ec-4410-94e6-6776a6cd901a · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:50:30.273755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T18:46:04.926179Z digest=sha256:0b08a0fa14ba816b21b978715f838c051690b94f2009c657f68654bd7e02c922

Observation 9777d95e-76ab-4c39-b98f-ae42929a6f48 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T23:21:42.110167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:21:42.110167Z digest=sha256:1f0ff760306b557d85b5e7ab2a80c7ad0d59a459916853da45ca10def4cb8f64

Observation 310bac23-4950-49a3-a4ed-0b616205c63b · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.415115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:fa4b2b91f6b74bc325c1d786790273fbca3f625db5f23ed5f97db8795f61f0ec

Observation f05640c3-6400-4fc2-b46b-aacd0f9b81f3 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:36:19.950190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:36:12.915133Z digest=sha256:d2f83dc1c5399e5f76f58d3ca4a1522261121dcc10abc6fdfdb033640235d3b2

Observation 03866696-993d-496c-93ff-04e990ad6929 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:30.196922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T07:29:14.545746Z digest=sha256:01a3422e9ed376cedab89d249b6e1d84ce148a97222248694a4378b5beab3ec4

Observation 4246816d-2079-4874-b669-88e413c3a234 · inbound

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices cites this paper.

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:50.166136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T07:57:49.746594Z digest=sha256:0614e801914b9caa0531163796a85872dec58291f9721683fad68eed5a4037fa

Observation bc7bccb3-9cea-42ea-8cce-3c2bf492175e · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.633191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:284734d9c306d075a20d8cc33775faad0379a447a2779af29fa76c1cb2ae323c

Observation 9e9c1558-3a9b-4696-8b17-2112027836db · inbound

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression cites this paper.

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:12:34.512202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T19:10:58.845006Z digest=sha256:65ce7205c6f082a4f7f7179dec9d2aac5a15f8b51f656df38873f272f7762651

Observation 07652548-2994-4d7f-9f53-23a64a978848 · inbound

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation cites this paper.

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.396687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:28:14.160341Z digest=sha256:865ee42e3b94e14a00d422988ff1ffd69dc973cb3f45e83d2814e6573ce7e28f

Observation c504d84a-0b2d-4614-a750-898e16cfb00e · inbound

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization cites this paper.

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.899516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T13:07:53.001390Z digest=sha256:6d345ca3b426f49c4968eb0b56e53d26d19e05802ba270a8cd84073ca7c5f9c4

Observation 66defc3c-1c96-409c-be9e-3130e323c914 · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:41.125011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:41.125011Z digest=sha256:5cf347deaea84d8797a5b0fbb24edb4f5c20ba8219c8006a79579d892e944505

Observation 5b0249f8-bee6-47b8-9021-6f42da1652e7 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.948303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.948303Z digest=sha256:a5423699ccf5a4751deebcaad6c140a356bf1390f456d5e4ec47e03902fc56ca