Pith. sign in

Paper Citation Record · LEDGER

PQCache: Product Quantization-based KVCache for Long Context LLM Inference

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2407.12820.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12820 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:29:13.944332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T08:12:01.964814Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc2cedf8-7723-493a-9b8e-dc54aff2add3 · inbound

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval cites this paper.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:01.969232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:7f1a13eff0595e85aa136199907994b0199ba214932adab0e309c98ce7813f11

Observation a0e7f2bf-fb87-479a-8ca3-3ca35c60e3a2 · inbound

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs cites this paper.

Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:13.944332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:13.944332Z digest=sha256:82511003e844a1afa19e1b68c38e4a09515b7efb9b67037225b477f499c519dc

Observation d06da1e1-b521-4fa2-8675-91a428ba9431 · inbound

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization cites this paper.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.391215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.391215Z digest=sha256:4b90006cad3b69a8cfaef53e29433cb02b15ffcec0de110870220822ccc93f7d

Observation f0d3eeb6-e310-4965-99c8-b85aa2d41096 · inbound

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling cites this paper.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.152541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.152541Z digest=sha256:13f84c0d400ebbf8e920e739b738797db90aca26a1f9098f6a7bcc57bac59be9

Observation 506d2d3c-2bcb-4aea-9efb-b6071374512a · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.745585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.745585Z digest=sha256:95d3ffcc1f1b01398dbe15c884c26d632e3ed5b08cf089d2cd0994acb5609e91

Observation 0a5774c3-cbdc-4a02-aab5-6c6eaea9326c · inbound

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding cites this paper.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:40.900312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:40.900312Z digest=sha256:ced4e6335a07bc56030c39d3542338253ccf0db3519d3d4f9f6ba244f2700ba3

Observation 2f441ba8-0dff-49a9-a7c4-17d4fbf2545c · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.220828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.220828Z digest=sha256:614c6e1e6b16ca9fa3e6ae6bf929e47686497058fc1d930f83586b05c6e092c6

Observation c88da4b3-afe2-46ae-b2bc-57b68f33f65d · inbound

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs cites this paper.

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:51:01.188085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:46:43.682254Z digest=sha256:c880bb06560f937bc50ee4bc8d19f9cdb1e5550224722eb60a31cd71824b89ec

Observation 58f9a7d8-cea5-49c0-9c7d-b171f4d489bf · inbound

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization cites this paper.

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:06:04.343444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:09:11.684432Z digest=sha256:ee114f112c6083e47eed7a7c512d7f4337e4738aaf338559c4627eb24050ca28

Observation e1473abf-d04b-484b-a419-926e0cb80548 · inbound

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding cites this paper.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.799232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.799232Z digest=sha256:472abb664119f9749edefba2d434755b9454ef1889e401bc8a3716d22bb44eb0