Pith. sign in

Paper Citation Record · LEDGER

D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2403.01876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.01876 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:15.072316Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T08:55:35.124659Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5fe69c5c-7bc4-4766-9cf6-10f572dab38b · inbound

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing cites this paper.

HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:44.201438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:43:44.201438Z digest=sha256:fc583984f24bc6df0980ef7aba9e7c1d0399593c3463deaddf27f257d2078447

Observation 5af68768-2a68-41cd-b293-fe6cf3b42cc4 · inbound

SYMPHONY: Improving Memory Management for LLM Inference Workloads cites this paper.

SYMPHONY: Improving Memory Management for LLM Inference Workloads D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:39:11.672916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:39:11.672916Z digest=sha256:feeaf21dfec3d54acd454a2fd3cd70248080c3686a816bbdfda919bc6c65e874

Observation c468697a-98fb-4f77-9ccc-01de47d26606 · inbound

Managed-Retention Memory: A New Class of Memory for the AI Era cites this paper.

Managed-Retention Memory: A New Class of Memory for the AI Era D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:54:35.875840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:54:35.875840Z digest=sha256:53ba56ff0be5bda6858ee597a3a4384b9e5c5fe92bfc911122cdbbb6d6523858

Observation 5ca33614-8a58-4420-bd69-e69bb42558f1 · inbound

KVDirect: Distributed Disaggregated LLM Inference cites this paper.

KVDirect: Distributed Disaggregated LLM Inference D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:14.667710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:14.667710Z digest=sha256:4fe045a398d2c4f9b37b60c9bf98a1357228eee2751614ebc79d4d4f73e82ab0

Observation 9a4ba963-276d-40f1-b6d4-18fd055878b6 · inbound

DeServe: Towards Affordable Offline LLM Inference via Decentralization cites this paper.

DeServe: Towards Affordable Offline LLM Inference via Decentralization D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:16:53.087337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:16:53.087337Z digest=sha256:a5510efe34bd3e4e510f573a8aeb411f7da574e5c28ea72fd00f49b006f7c637

Observation 662abd6e-5878-4a70-b9fb-c2d4c8dd6cf4 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.045273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.045273Z digest=sha256:d32badfd1d8161512228704bf915324f40173fb9abe4bcff50fcc3000561005f

Observation fb9321c0-4d56-464b-bfc7-346b861a505f · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:39.355105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:39.355105Z digest=sha256:dabec3746973709972cf40c4913ac2dcbc18b946d7f5499c0c92541fa456da5c

Observation dbf355d0-cc15-4c0a-bf76-509991c09e31 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:45.318574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:3f8bbbd6a4517c752cb748fd44eca42e491534bdb8c8a8035b1e0a8fceb0eddb

Observation e1bd4be9-c9c3-4191-96a6-6620016ead29 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.126148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:80107c3ed7e0de1bbaff4ea7585fe9d7dbacd87b1adca0a759d25382959cd1ab

Observation a99dd34b-bc03-4b6b-826d-0957edd36c35 · inbound

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack cites this paper.

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T20:58:04.182764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:58:04.182764Z digest=sha256:8a98a24f23dcafa325e3eda08ac92f6527ef037227cfc889baedcc51ff14d99f

Observation 60fe65b5-c7d9-437f-aa3f-c6cb4b316c79 · inbound

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure cites this paper.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.357426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.357426Z digest=sha256:b512c31bd3c3ad5bd762b920b174c1a3cd201f50fb4595e775e111c569ba7402

Observation d67a3e2d-0afa-42a1-94a8-e409aa77b131 · inbound

Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System cites this paper.

Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:15.072316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:15.072316Z digest=sha256:00c6881246b5fec3088793ba6533631f1ad0c404fad78d15b7fb3011bfb2fb03