Pith. sign in

Paper Citation Record · LEDGER

CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2310.07240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.07240 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:00:04.086265Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:36:44.087990Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 78ac79e8-f742-4930-865d-5fec79050c14 · inbound

LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts cites this paper.

LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:00:04.086265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:00:04.086265Z digest=sha256:73a34253a9fd262241219ac07bd1f514b8c4247ea0b6c46a1f2c2e36ec7ea65d

Observation febecb8e-2517-4865-896c-d5737ddc14a5 · inbound

KVDirect: Distributed Disaggregated LLM Inference cites this paper.

KVDirect: Distributed Disaggregated LLM Inference CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:14.604922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:14.604922Z digest=sha256:72589e9b89f27040340c3b5ebc5b0b55c716aa11f76ae1f99038ea88968847da

Observation 12d1fef0-7650-4007-95a7-1d84b3e70cb5 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.739768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.739768Z digest=sha256:451de6024cc3d84660df1e34ccbb97bad60001bcd05944f73523a8e08ae4507c

Observation e560a3a0-8f46-49e1-8f74-80956214fa5a · inbound

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation cites this paper.

Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T05:40:21.554321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:40:21.554321Z digest=sha256:dcbcb8e36100848002b18fe2691c653f8a553024448172417c67d505e8e6ce12

Observation d04da29a-e91d-4121-9b85-9c9c9a1bacee · inbound

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding cites this paper.

SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:44.937596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:15:44.937596Z digest=sha256:1db977ef1845e62cd0999138e1e6276f061706d495b4586d680edd515517ef41

Observation 8d66700f-ad39-4d58-a4f0-4ec86695dbfc · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:26.441620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:26.441620Z digest=sha256:bc815fe7dfa9198fdf360e394beefcee45ad7f0b0142f20217aaa4b253a3b529

Observation aa5dc71d-861b-48c0-9a3a-05f55e5e1dc4 · inbound

On Evaluating Performance of LLM Inference Serving Systems cites this paper.

On Evaluating Performance of LLM Inference Serving Systems CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:15.906010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.906010Z digest=sha256:4bd953403080e275245d2b4a5031f566c3a5a18dad49e862dfdd826fbbfca474

Observation 347e9ca1-f020-4eed-8969-5d0979cfd26f · inbound

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems cites this paper.

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:28:37.932904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:28:37.932904Z digest=sha256:5842e1f3814c20f99ea636beaf6ff0e7c37beb1ab3dcb1b6b6f6e11f32ab9b93

Observation f388e6c0-2d7e-45fb-a46c-9d99ab23b9cf · inbound

Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution cites this paper.

Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:42:45.723763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:42:45.723763Z digest=sha256:54678530a3cc35011dd48be76160e9037088a20ec8044c8b849d318467411c6a

Observation 09e55922-d738-414a-9c2b-16ce9c3cedde · inbound

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving cites this paper.

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.970739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T17:23:13.154458Z digest=sha256:cb3e44955229ade47659eb31bf116d401bb0e40a179e73dcc13605b5a70f7333

Observation 2694c557-0caa-45ea-8992-a1bfed5e6542 · inbound

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving cites this paper.

QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:57.479670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T01:47:43.240850Z digest=sha256:13c3c84d41c7ee2591fd512069c9a50d11254fae6905782cda5f3c62073ec6d2

Observation e221043e-7543-4001-a33a-60b3ebba5179 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:45.315356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:5ce5406a9a7b5b5f55975e82e551737a44eea71436ba275e7dcb8b911e5060e2

Observation b66cc25a-7a9f-4cbf-b764-20fd0aedde85 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.115489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:28ab2b5e1dca28db9ef5ff0e814d3733f1266cfc94f835c0c84c8ab7c5f64858

Observation 0a8d3eba-7310-4e24-8daf-309770ebc037 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.089166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:d7036d082a362aa798e8a4fb7985eeb5db08c13e702a3e1b80d27e8d63e0a4b0

Observation dcc0f714-5a6a-48f1-8a67-3630cf0085b9 · inbound

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel cites this paper.

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T02:10:49.534551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:10:49.534551Z digest=sha256:259c50e4226f34cf1423af919c956b3df685c9fdd921566cbd80cf1cf9b75e3e

Observation 2de3a19a-a0f9-4dc1-9201-8d2d0e35ce6d · inbound

Persistent Computational State: A Session-Centric Runtime for Generative World Models cites this paper.

Persistent Computational State: A Session-Centric Runtime for Generative World Models CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:26.861103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:46:26.861103Z digest=sha256:106753e33ddac6435be4aa9278cc9f23583c5f702281c1d476c251c248944da8

Observation d8f82b89-0ba8-445c-93ef-2f2041318699 · inbound

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework cites this paper.

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:33:27.633007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:33:27.633007Z digest=sha256:020ccd7ad443dd2e179ec863a8ce8ccd37437a95a4ec8cd922e9f08055d7202f