Pith. sign in

Paper Citation Record · LEDGER

CommVQ: Commutative Vector Quantization for KV Cache Compression

As of 20 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 9 inbound Pith citation observations for arXiv:2506.18879.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18879 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:49:24.164784Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:54:14.045740Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:57:26.279666Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b18c4fc-0c2f-4622-a726-81a34f765043 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

CommVQ: Commutative Vector Quantization for KV Cache Compression Categorical Reparameterization with Gumbel-Softmax

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.080945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.080945Z digest=sha256:e80e8265c5ce92309ac62957545200dcb060af8f0fc46797897e766c47637cb7

Observation 48bb35fc-d9c9-48e9-b944-3af0ca493d08 · outbound

This paper cites We calculate the codebook size based on LLaMA-3.1-8B-Instruct model.

CommVQ: Commutative Vector Quantization for KV Cache Compression We calculate the codebook size based on LLaMA-3.1-8B-Instruct model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:49:24.727617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:49:24.152304Z digest=sha256:53431235b3614b3305acec6efd9c2955b257bbc9c60383731bab82b0ae96fc65

Observation b7e75f07-be45-4b0d-846a-5363ee659409 · outbound

This paper cites Residual vector quantization for KV cache compression in large language model.

CommVQ: Commutative Vector Quantization for KV Cache Compression Residual vector quantization for KV cache compression in large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.098559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.098559Z digest=sha256:e02e4b92b531be8de9d9d27161a73e96c0fb799f45f452602aef8c1e26d8cc12

Observation 80dd43e2-f99f-4afb-a5f6-826162e403e7 · outbound

This paper cites Crafting papers on machine learning.

CommVQ: Commutative Vector Quantization for KV Cache Compression Crafting papers on machine learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.103731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.103731Z digest=sha256:8a7f5de50e338e2d495ba3f439677f7fdcbf8d9d2837f0064c603f0add1d3dbc

Observation 26ac9b3f-8c7d-44fa-a441-d8bdc8ef82f1 · outbound

This paper cites The MSE is evaluated on a small subset of the FineWeb-Edu dataset (Lozhkov et al., 2024).

CommVQ: Commutative Vector Quantization for KV Cache Compression The MSE is evaluated on a small subset of the FineWeb-Edu dataset (Lozhkov et al., 2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:49:24.691393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:49:24.164784Z digest=sha256:42178beadbb316d34b83f20fc1426f2735bde0bca8393e427114d118c95a55fe

Observation 93a778ed-c57c-4e7a-ad1c-0000ebe9a475 · outbound

This paper cites Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption.

CommVQ: Commutative Vector Quantization for KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.114689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.114689Z digest=sha256:f3d887a65f9e9f0dfbd8e46190e1970ff258caeb1fd50c33055f36ccf55cfbb9

Observation 15f6826f-8272-447e-a03c-6f8e46aabf91 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CommVQ: Commutative Vector Quantization for KV Cache Compression Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.121353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.121353Z digest=sha256:dbc9c08b00362fd685aa4535a5064817db40269caa32e7ed265474b76360fde2

Observation 3c0e767b-8816-47d0-b0f4-7ad702fe31be · outbound

This paper cites Qwen2.5 Technical Report.

CommVQ: Commutative Vector Quantization for KV Cache Compression Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.133333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.133333Z digest=sha256:363fdd9da7e3d46852ea45a19abd702f4781d208dca576b5c547bd00b116c583

Observation 361296f1-a515-4717-a4d0-4c83f08c7bc0 · outbound

This paper cites KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches.

CommVQ: Commutative Vector Quantization for KV Cache Compression KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.139988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.139988Z digest=sha256:5f2c74fcb172831a1493981e7743a482781159d3f1d18049989dcdcab142ccc7

Observation d50bfacd-1de6-423b-801d-20c0c0420fe7 · outbound

This paper cites KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization.

CommVQ: Commutative Vector Quantization for KV Cache Compression KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.145639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.145639Z digest=sha256:bf680935d40d828b4b8e0fa065b314d047053d5d18a2c17e5620d5244faceedf

Observation 18f02273-0b23-4e18-b4de-5abc9912a1b5 · outbound

This paper cites an unresolved cited work.

CommVQ: Commutative Vector Quantization for KV Cache Compression Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:49:24.710831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:49:24.159283Z digest=sha256:9a8080f5b13d4042e47e1743d2a79504cc5446be02d3f842c992fe7cffc72b6e

Observation 2baeb153-49e3-4b37-908b-b498371ecb17 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

CommVQ: Commutative Vector Quantization for KV Cache Compression KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 1984

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.074952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.074952Z digest=sha256:447ad88e77da2717ab21d6f0529d67ec1c280fe27b1a492cf2c1c57426f1f4ec

Observation 8352d7d2-fe78-45cd-8d46-28f321f67305 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CommVQ: Commutative Vector Quantization for KV Cache Compression Training Verifiers to Solve Math Word Problems

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.057637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.057637Z digest=sha256:4ebdab10d6164f8fcac959a6c5c33632c3d8735022ceccfe39d6aa8ed3366593

Observation 9548ebe1-9d30-4043-9d94-17531f1a2ea7 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

CommVQ: Commutative Vector Quantization for KV Cache Compression KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.109429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.109429Z digest=sha256:7d71840a5c7b4950bbe4c2a63bd2e81702ad133e3c17375420b5ca204fe5ae03

Observation 6cec6971-1476-4399-82f1-8bd5e94784c0 · outbound

This paper cites Mistral 7B.

CommVQ: Commutative Vector Quantization for KV Cache Compression Mistral 7B

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.086931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.086931Z digest=sha256:f004753af4226a2b6fb59da4e0dc301882455fb7752a207772e3d8233183ee6b

Observation c2cfbc77-1ff0-4008-90f5-5a76d93442ce · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

CommVQ: Commutative Vector Quantization for KV Cache Compression LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.050006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.050006Z digest=sha256:9abb5f889bfb5d047ee05a9ab4d22c0d2162708fe525deb9c2441624e61d0d3d

Observation 8d097a56-f1c6-4501-be3a-ecd5698b98ed · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

CommVQ: Commutative Vector Quantization for KV Cache Compression LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.063087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.063087Z digest=sha256:eaac21e9076de04cc57345dbfd6b14a1cf3829bcb3c4b97c02d2aa6c838cdb8c

Observation 16f230de-5f9e-4ad8-8d97-b83fc7f1ba81 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

CommVQ: Commutative Vector Quantization for KV Cache Compression DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.127218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.127218Z digest=sha256:08e9bf61a98ff31013f9fea64fab77e24f9b3dd0ae15d5a0b5305b07982db172

Observation b960514f-b1aa-4abf-b6eb-716655140933 · outbound

This paper cites LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning.

CommVQ: Commutative Vector Quantization for KV Cache Compression LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.092793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.092793Z digest=sha256:e1fd45117783af437be06d218a12e992c9cc20a99655626992e92ff6effb9cbd

Observation 830942ce-3018-40fe-bfde-9532c3878087 · outbound

This paper cites The Llama 3 Herd of Models.

CommVQ: Commutative Vector Quantization for KV Cache Compression The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.068972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.068972Z digest=sha256:c22c8d8db9ec71f80659d3c43db3d40eac2057758ab6ca9f2c778de8576d31da

Pith citing papers

Observation 76071d9d-43e9-4def-a737-eae415b75c2e · inbound

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization cites this paper.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:00:44.674074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T07:58:24.456859Z digest=sha256:4a688c7276df4c3bbfe0f332aee81f84c71370f510d9d5044554541259b61944

Observation 5a97a9e5-6480-4ca3-863f-8a8defbb5ceb · inbound

Online Vector Quantized Attention cites this paper.

Online Vector Quantized Attention CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:44:11.453481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T13:42:51.275943Z digest=sha256:8cc08d74b9e1344256dab30d289228f7b299464a02ba6ce435e957c0a104da56

Observation b2271486-0af3-487c-b0b0-404416548090 · inbound

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving cites this paper.

SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T03:14:08.310691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T03:09:51.453839Z digest=sha256:237e1b8bb9c7ccc54ba1a9159b29a11001c54b07f2e1f18d25086b8e5ff506ca

Observation a9c3e704-b1b5-4523-9399-0a99bbd6c1b4 · inbound

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture cites this paper.

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:44:45.580150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T14:35:02.484377Z digest=sha256:dcb598ea3076b10a9dfe617585e9c109c98e7dd35f35d04d38cb3a37f68dcbb1

Observation 6a04cad5-0cca-4480-9555-201c7495f323 · inbound

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control cites this paper.

STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.281085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T18:34:20.645677Z digest=sha256:ecbd61b1f49a3c9a442dbec7bedf955cd34100304d15afe0cac17efb59eb4ac0

Observation 96fdaf1c-f2fb-4cec-adb7-d8c04a040b56 · inbound

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference cites this paper.

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:38.209120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:38.209120Z digest=sha256:ac8a8a0820d9ba38ecd62f5a39bc4a31ae69441f9967fd7d23cee612c7423244

Observation b218b01b-f5a8-4ec9-bc57-da5b4c7cf1d5 · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:50.064347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:50.064347Z digest=sha256:160074010a0c40dbc97d3d5953fce9dc66f0cdaa039931f79bfa3db77e4c8783

Observation 8e17b076-902a-442c-9477-b77ee178bbda · inbound

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms cites this paper.

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:14.045740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:54:14.045740Z digest=sha256:8610d3b5d608950bf7bce72d868d5d3c7b69ec0fc6199082df5ae5b077d63743

Observation 3b5d9bf4-acf6-4453-bb33-331adf66bcd4 · inbound

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation cites this paper.

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation CommVQ: Commutative Vector Quantization for KV Cache Compression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T12:54:26.530296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:54:26.530296Z digest=sha256:9a50004efba38e7930aea0e6c5f095f00cf9659e20fd4076da9ef27387429fb8