Pith. sign in

Paper Citation Record · LEDGER

LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2410.00428.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00428 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:10.207787Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:24:56.959321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2e5256b4-3516-4ab0-a61b-4a7441e3b7f2 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:49.572092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:49.572092Z digest=sha256:d6b009747c0c2fc64b3f4e968d5a9a5f6ca9f9f0f989140c92b0342f56c09913

Observation 702c9932-288c-49ce-8502-41fabbff8863 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.098928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.098928Z digest=sha256:0f14fb6195cd8454c0722371d634b9033515d39cead32966ae72410a78d8d4d4

Observation 01bc516e-8c62-42f4-bf50-12890b7a941e · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.207787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.207787Z digest=sha256:71e3d1dba3415670279a9822ea86a533746c7d5b9fcd3a778acbb2ebed4b4389

Observation f770c5da-57e3-4456-a076-a60800a339bf · inbound

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration cites this paper.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.359687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.359687Z digest=sha256:5d9ec2bdd3615d4f7ad9a6ff0892cc2c3dbf75d85fb1c78b0ce13cf0d2cf0d5e

Observation bca2aa03-58fd-45a2-adcc-31e71317e029 · inbound

Efficient Remote KV Cache Reuse with GPU-native Video Codec cites this paper.

Efficient Remote KV Cache Reuse with GPU-native Video Codec LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.824323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:543ee5a7b4737ee05eb6d990b37782cc98cee91d558495ec0b98a0586c631feb

Observation 0b7c41ff-21a0-42bd-b047-83131df9db7a · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.903875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:c3eece24ac3a494b034a222b1129a1ed70d6230a437d457f42fb4210723d9391

Observation 92e32207-aca9-4950-81da-403be932bc1d · inbound

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving cites this paper.

Adaptive KV Cache Reuse for Fast Long-Context LLM Serving LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.960758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T17:23:13.154458Z digest=sha256:d4c5b2fa0604441a5ed3724cf75753aae1feff82e8c521b13370527a4bf88b77

Observation baea7759-e46b-49ec-9f53-6230a91772e4 · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:10.825037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:10.825037Z digest=sha256:0f27e5217f17eaf97c7660bab5fc6e0cb26d6e126130e0d5f18925dbf3ff5b16

Observation 2a1b2494-8811-4498-b14e-4b0f6b4fe1db · inbound

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems cites this paper.

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 230

Resolution
unresolved
no resolver link, observed 2026-08-01T19:51:22.631756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:51:22.631756Z digest=sha256:6587698a427965fb486cabbdd9f9418fb90f9b72fabbf7b2276ddc82f4ca2eaa

Observation d91cb980-1b46-4f98-b2e4-c768f8990f3f · inbound

Persistent Computational State: A Session-Centric Runtime for Generative World Models cites this paper.

Persistent Computational State: A Session-Centric Runtime for Generative World Models LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:27.035461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:46:27.035461Z digest=sha256:e7b99c99f5bcff0be81d68c7d09136f9ae1adea0616c1531fe91ea5e841b1fd3