Pith. sign in

Paper Citation Record · LEDGER

SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2306.03078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.03078 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:24:19.403461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bb0773e2-c299-4c5c-b311-4e488060f525 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 196

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.210769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:1c86f932fd9d8b0b5f567ff59249ee567215284ea45bb8c2532fa1aa27d3b7bf

Observation abbca5a6-faa9-4c66-9b1e-0a241ffeff3a · inbound

EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices cites this paper.

EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T16:34:59.286285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T16:34:37.083239Z digest=sha256:87898a83c6aaf8475d86dba93c546409029351cbb81097489ae0ef03419ee0b6

Observation c47f2b99-b8d2-47f0-a23a-a326b70c7a08 · inbound

From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction cites this paper.

From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:47:07.687089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T05:45:46.354637Z digest=sha256:3f4ba21aeddc9b60764156044bf995e4743f7a8b889b34071ac2724f78b8b1a2

Observation dc117cc0-8f7c-4e3e-94a1-dbcfbac4d388 · inbound

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM cites this paper.

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:24:19.403461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:24:19.403461Z digest=sha256:3b4cc8fbb7939da4f009b419cef75f41c1e44d1018f3453fe66fcf4aa55ebe51

Observation 8c40f42f-b3c2-42b6-ac17-82841a3ccebe · inbound

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices cites this paper.

Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:04:34.644179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:04:34.644179Z digest=sha256:c791529d444ec7ff7ac6d5e36194d8d1e58dc27510f110281542a1c725ebba65

Observation 5199e9da-5a26-450e-8ec1-9e215102956d · inbound

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference cites this paper.

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:26.060146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:26.060146Z digest=sha256:acf9c0132ddcfdc19566efb68a167445b78b4e91a198d9868a340a0f6f40fc81

Observation 56105472-385e-4698-963b-1c5ff359324d · inbound

Fast Entropy Decoding for Sparse MVM on GPUs cites this paper.

Fast Entropy Decoding for Sparse MVM on GPUs SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T19:38:56.284753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:38:56.284753Z digest=sha256:4e6138bad07ede90ed74eecbb70ae8784e7bd4ae10a17c4115c8db3198813b0a

Observation 912358a8-8cba-4513-9622-c2cc17e05a8c · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 124

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:a713a37b07394503e49b973099800e1084c9982a04240d717a222950d9fb5008

Observation d2d37f02-7cb0-4f3a-aed4-d8319b61500f · inbound

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling cites this paper.

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:56.668692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:10:06.994557Z digest=sha256:d6d9becdfa1495abaa0df4bb85a4bb40936a5d044ba16dd35cc65a5441e119d2

Observation 23fd9b77-2b6d-4bb9-8ccc-76ff4ee2878d · inbound

TStore: Rethinking AI Model Hub with Tensor-Centric Compression cites this paper.

TStore: Rethinking AI Model Hub with Tensor-Centric Compression SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.081572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:16:58.389670Z digest=sha256:345b431010f4a6a946d6bbc07aa7f7e966f7692bf2d43c8a6ac0b9bc4087ffaa

Observation c9f6c509-6706-40b0-bc7c-e2b5e87c82f6 · inbound

TStore: Rethinking AI Model Hub with Tensor-Centric Compression cites this paper.

TStore: Rethinking AI Model Hub with Tensor-Centric Compression SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.731475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T22:04:16.100009Z digest=sha256:e0bcd3d5af7aa611c3a544f04ddfb3ba6df78185d29bdcfd905c9f523040854a

Observation c7d73a95-d2fe-4d80-b28c-16bb1f9a002a · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.324016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:29:51.182114Z digest=sha256:6e2c3ccbb595ff0c84450a569e69593d6f9e38c6d35f99510063b14334a8946b

Observation 137de9a0-079b-43b2-b6ed-fab5d6b3da09 · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:02:42.165210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T18:01:08.514022Z digest=sha256:04e9081a3a169bc4795bf151cc75c72d7d73a394f3a2a8bf37a0181f39431ec5

Observation 6e088adb-7b8d-44b9-b8f5-fa52471cfced · inbound

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation cites this paper.

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:04.928422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T03:04:14.900791Z digest=sha256:81be96f352fa9acc8d8cfb592b469686f5443627adc0d9876a350d951bb4bd4a

Observation dd2d1e82-23ad-468d-8bf9-35cbc370a7c9 · inbound

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization cites this paper.

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:38:17.222652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T02:32:50.182859Z digest=sha256:127ddf39315f82c3cf8f13e8af0964137288e9ec6fc16e5fa8c45cda137911c8

Observation 402acaad-10c2-44e7-aeb3-3b5d57c70dba · inbound

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales cites this paper.

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.489349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:44:42.989053Z digest=sha256:c23612a5ac44809f6700a55143d9eea4b1195a2731a8626ba277c11184325054

Observation 6c94e8d7-3868-43ef-a1ff-2d9524769418 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.423526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:38ce06025b5860fa60c06f5bf4c2022f7a50e7a5ce94b1440a40e26915f29b9f

Observation 04fc6925-320a-4bf0-9578-6fb8c4805269 · inbound

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels cites this paper.

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:16.276376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:35:02.120009Z digest=sha256:711515be0a9d1b947f1bca4c09908a04011dac89267642902e7a00e68e47aaea

Observation 508720db-69cb-483d-9c1a-ca2eb750e94e · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:42.449581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:f315ebce457972f04d479948cedd8e299c3c4e5ec2c93b4e3d4bf92c5bb5bf56

Observation a31d2bd8-75ae-4446-84fa-f321c3bdd100 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.092490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:8cccad8f92b9f742c7afb6230b71785bf93779ca509d058d72bc642bdbd8fb77

Observation 7c4972d2-18c2-42e6-84af-5789017e334a · inbound

Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven cites this paper.

Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:08.487611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:07:51.451501Z digest=sha256:7057b24e593bc1c01355892c59666814b8f04831693909daa9c8f6a6f69d5e08

Observation 1bb94b70-b3d9-4bcb-9d5b-38cb1b382a7c · inbound

XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA cites this paper.

XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.301572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:13:29.486007Z digest=sha256:fd722e2ed8c243148a228b4fed4c32d1b90daf6577aab398b6a7e799613cdea4

Observation 69fc575c-709f-4493-856f-b0fbb3ae6b16 · inbound

ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models cites this paper.

ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:37:07.849512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T02:35:25.946090Z digest=sha256:62ef84e0780eeaf59de1e53556da7790328e63de12ffeea323bda281ff2f9f1e

Observation 25a9cd37-c230-48c1-898e-23fd8e314672 · inbound

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference cites this paper.

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.693173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T21:28:36.358474Z digest=sha256:ac97e18fd19e229814d63f25787d3b0b5605857b8a8735458e045afc28897e88

Observation 525299f0-dde1-4505-bead-c2e73308cb40 · inbound

StatQAT: Statistical Quantizer Optimization for Deep Networks cites this paper.

StatQAT: Statistical Quantizer Optimization for Deep Networks SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:22:55.766489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T01:20:06.033991Z digest=sha256:2aef3ac389f58e96fd68a80ec17d9aa954b9be2a6b03eeb6fca009d1cbad4a35

Observation 7f9f3ad1-5da1-465e-8b40-aa5a0d1a5d26 · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.685044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:714b31e70839015398572964489afafed054e2876df27794d23126301a3f1ba1

Observation d74e7e15-7d25-4983-b79c-95b16e64e007 · inbound

Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification cites this paper.

Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:39:57.649753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T09:35:10.701255Z digest=sha256:b7eb4350016d104d65efc364799abfbd870e8fd6e095c95529a5f281faa714c6

Observation 6785d84f-8a65-4c56-840b-839f7360b804 · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.947415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:412c40049c43b26b8d941e582cfde215150ca31dd93693ed596a0ce3716712ae

Observation 1024fff2-658c-4bfd-af25-aaae7e16e62a · inbound

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models cites this paper.

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:47.964211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:54:56.386488Z digest=sha256:6c867a413ecbb8238c2f4c3c5fe8e415098cf08860f004bdd4d399d22b3cebc6

Observation 502cb8de-5919-4714-b595-88ad5a0b764c · inbound

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation cites this paper.

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.443665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:28:14.160341Z digest=sha256:ae8073079b6b70b2f918e1a49980c12fd43cac46942775d0cce05f6d515ffc8c

Observation 717bb3ab-73a8-48f7-8b9e-9643b389c5e3 · inbound

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning cites this paper.

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.563117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T11:27:02.902720Z digest=sha256:4b91904633b841561b06883fbc59e360c74a8e9e979e4d9b14b6901a9871b572

Observation 3ee4bc38-d30c-49f2-ad35-d4906aae8c83 · inbound

Do Transformers Need Three Projections? Systematic Study of QKV Variants cites this paper.

Do Transformers Need Three Projections? Systematic Study of QKV Variants SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:36:17.239860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T15:14:49.475561Z digest=sha256:8a6468edef8c1aa996c2e9e4aca76da763ffed060d66da64a52b81eae7101cf9

Observation 1e05ef26-205f-4fbf-bf74-2ac4256e16f8 · inbound

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models cites this paper.

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:56:44.617007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T07:16:45.665440Z digest=sha256:3f5841954878699428bf64581d9e7bf1e719248c8ab03faaf9463a5c39a1e97a

Observation f96606f8-32f2-469f-be75-f20943798333 · inbound

ReCache: Learning Budget-Aware Caching Schedules for Diffusion Models via REINFORCE cites this paper.

ReCache: Learning Budget-Aware Caching Schedules for Diffusion Models via REINFORCE SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.437227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T02:09:44.280357Z digest=sha256:cb4fa9e2e55a72074c67d7610ba7dd05fc11d157e2deddee6d61fb96bc0f1644

Observation 4c7493ab-d888-45ef-b06e-f4f834f5e3da · inbound

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation cites this paper.

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:16:15.998711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T15:37:34.129339Z digest=sha256:89abe22e64aba1e8d131b56844e137e308ba919e66f93d1233b178c9d77dbaca

Observation fc55f228-6f35-4346-95aa-4801b5f1abe8 · inbound

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization cites this paper.

TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:48:20.625511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T07:37:59.122704Z digest=sha256:25cc654db998705c388e2863ceb0d3d54b5fc8daa7f0aa402c1d0d61d0eee300

Observation e9202480-a0cc-4eb9-b907-2e86d001c1fb · inbound

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models cites this paper.

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.520230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T08:46:01.500880Z digest=sha256:d02f6acc4329abacbb7691bdc83b5b1f83bfaadf807a4c23840a14a9b11e9485

Observation 0be784fa-4d15-4337-ba3b-dc65e0879361 · inbound

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers cites this paper.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.414192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T14:56:10.553212Z digest=sha256:04c94cf325a502ca84c48f51eeddb1a31c9b01602a3c23a85546e76190863e81

Observation 1608e6ca-b4d6-4e66-970f-76d12497b53a · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ce95d6f0ebe34cf454f4e28af4f7b83e56be5e8200af10df84b4a4c92567a5f9

Observation 985908c5-d4d4-4e43-a1c3-239b203a1eea · inbound

Reliability Scaling Laws for Quantized Large Language Models cites this paper.

Reliability Scaling Laws for Quantized Large Language Models SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-14T08:45:52.855783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:45:52.855783Z digest=sha256:d0c86dbab24808bb2444024d7c703e6617811deb7720e775f5d8bcd116d6d06d

Observation c93a4183-b56e-4ef5-ab2f-0cba141bdf49 · inbound

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems cites this paper.

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:55:40.314783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:55:40.314783Z digest=sha256:23b11a0e5db68df6124fd49b3b67daf0a719462ad29cfe9bd496baa2465759e3

Observation 2c55b6ea-52d1-422e-b93f-4f07719a5c3a · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:06.957974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:06.957974Z digest=sha256:5b16c1e76c3e89e917ea0d065df37153e692961b49579408a0a6a8f70708aa1e