Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference

As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2505.21919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21919 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:45.314876Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 375ff02c-a8a4-408b-b728-6828018802f6 · outbound

This paper cites CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.290344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.081676Z digest=sha256:50efdcd5625fa214a444ad665991ade7b98e1a215b67adfd0d0fda953fc08666

Observation 0799148e-4176-41c6-be15-11073ca27ef8 · outbound

This paper cites Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.181456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.174156Z digest=sha256:065c37021271dd9a8f5a9cd0d4cc50a229e039b0efb78fb57d08f0c3fcad9ce9

Observation f05966c6-ad4a-4800-b5df-a7844b415cc3 · outbound

This paper cites Unlocking Longer Generation with Key-Value Cache Quantization.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Unlocking Longer Generation with Key-Value Cache Quantization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.010693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.282406Z digest=sha256:de01dc40988eae31191f5c05ee567c2d3e498503cfa6601a9e4c4baaaca55006

Observation fc9c2d9e-0206-48e6-a911-c3301b27228b · outbound

This paper cites vLLM vs TensorRT- LLM 12, Automatic Prefix Caching.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference vLLM vs TensorRT- LLM 12, Automatic Prefix Caching

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.914847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.386461Z digest=sha256:dcd1bc45b096efd82e75764daab59c4a50d83e08dcf600e443e103c0a8a57ef0

Observation 9f28759a-2cf3-4dd8-8106-86b6acf5f8e9 · outbound

This paper cites More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.472480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.472480Z digest=sha256:67d67464ea0321410550413fd1d1c521726e543acdaa290dd5e93229ca219c2f

Observation ba612c96-afe4-4af9-9450-4a4d25d12738 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.554910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.554910Z digest=sha256:e63876c22d63b60895028be39bfff0e9b5cbc09bff82d85d966d01482bdf2a56

Observation 6f72440f-ebc4-4f90-b435-07f7a5963d91 · outbound

This paper cites Longformer: The long- document transformer,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Longformer: The long- document transformer,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.729746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.706234Z digest=sha256:37c9c6096ee50e2cac3ec5166297641aed77754931727cede7f72bf8db96252d

Observation d6e904c3-9c8b-42a7-95ff-0b9f54fb0e8d · outbound

This paper cites Pie: Pooling CPU Memory for LLM Inference.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Pie: Pooling CPU Memory for LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.784746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.784746Z digest=sha256:19c4cbedd1f40b94fb6c547552b852338310480e61bc0e99c084049a9d5dd682

Observation a97db814-edd7-46eb-84ce-9e6adf73d081 · outbound

This paper cites Mooncake: Trading More Storage for Less Computation—A KVCache-centric Architecture for Serving LLM Chatbot,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Mooncake: Trading More Storage for Less Computation—A KVCache-centric Architecture for Serving LLM Chatbot,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.548558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.875176Z digest=sha256:18cc8d05cefa098d0ee9d0889ea1a6c82ae7f505fa6a518fff7ac2e5af490c8e

Observation 0fb4072f-0a8b-49d0-9160-848e7435ae17 · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.334911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:44.988250Z digest=sha256:1227726b1612ebc7948358d01761b0ac24d3bddc90b6074ba68a44edce1c7efc

Observation 5aeb4ef5-c030-4ff3-a977-6aa2a4c68dd7 · outbound

This paper cites DeepSeek 3FS.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference DeepSeek 3FS

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.087572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:45.106737Z digest=sha256:7447bf89edeab7a58f60bc961f7229c29349331fb760e0bb40dc598f86f3c3a6

Observation 75256a89-6565-4e20-b516-4c4f2b4578ab · outbound

This paper cites IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:45.827026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:45.218533Z digest=sha256:6c090f9bbb4b19699f386c37305b7635c112df7e45409fbe2700e54401d4cd62

Observation e12c1413-4397-4f13-abee-b623222209de · outbound

This paper cites Exploring cxl-based kv cache storage for llm serving,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Exploring cxl-based kv cache storage for llm serving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:45.549127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:22:45.314876Z digest=sha256:f6c1dc2eaf920be2551a8ca06113cb15dcdc1608512995798e9dc8a73ebef438

Pith citing papers

No inbound Pith citation observations are available.