Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2505.21919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21919 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:45.314876Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 375ff02c-a8a4-408b-b728-6828018802f6 · outbound

This paper cites CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.290344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.081676Z digest=sha256:5ca6d1f4988423c8def2942cf236bed856873322f89eef8ebf10a6b4fb2f2e1d

Observation 0799148e-4176-41c6-be15-11073ca27ef8 · outbound

This paper cites Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.181456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.174156Z digest=sha256:fff6466f3967daa8116dc80f97b6eec27129f370ba78ae77c25dbe58b5cf722e

Observation f05966c6-ad4a-4800-b5df-a7844b415cc3 · outbound

This paper cites Unlocking Longer Generation with Key-Value Cache Quantization.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Unlocking Longer Generation with Key-Value Cache Quantization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:47.010693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.282406Z digest=sha256:2555533326caaa4804815464d1fb60684708778703c195afb3671d6cfcc4e7a1

Observation fc9c2d9e-0206-48e6-a911-c3301b27228b · outbound

This paper cites vLLM vs TensorRT- LLM 12, Automatic Prefix Caching.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference vLLM vs TensorRT- LLM 12, Automatic Prefix Caching

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.914847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.386461Z digest=sha256:ccb5a9da859e5a6059a48de38fee57cdc5dfce41cd83a1e8ceafa040fa53f575

Observation 9f28759a-2cf3-4dd8-8106-86b6acf5f8e9 · outbound

This paper cites More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.472480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.472480Z digest=sha256:8e421038c086b0c3bc7d891a9fb60e4a9ff4ad032e190e97117f2b6503dca800

Observation ba612c96-afe4-4af9-9450-4a4d25d12738 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.554910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.554910Z digest=sha256:79e51ac6a560534717ff270485c71c91cc9f17fac22f38cab79763d7f38392be

Observation 6f72440f-ebc4-4f90-b435-07f7a5963d91 · outbound

This paper cites Longformer: The long- document transformer,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Longformer: The long- document transformer,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.729746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.706234Z digest=sha256:253039f6ab9d574b0039833eb9a63738558608563ea9200fd7975ddfbb76a889

Observation d6e904c3-9c8b-42a7-95ff-0b9f54fb0e8d · outbound

This paper cites Pie: Pooling CPU Memory for LLM Inference.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Pie: Pooling CPU Memory for LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:44.784746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:44.784746Z digest=sha256:38756b4b7d4f2f6178c8760f1349c80d80e9697a833dad92ce6f52b3aac36439

Observation a97db814-edd7-46eb-84ce-9e6adf73d081 · outbound

This paper cites Mooncake: Trading More Storage for Less Computation—A KVCache-centric Architecture for Serving LLM Chatbot,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Mooncake: Trading More Storage for Less Computation—A KVCache-centric Architecture for Serving LLM Chatbot,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.548558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.875176Z digest=sha256:6f790865a06f6cc861db3ccfe770a0c38b95649b93a91f9843be046d494664f4

Observation 0fb4072f-0a8b-49d0-9160-848e7435ae17 · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.334911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:44.988250Z digest=sha256:7f718fda623375d1d9eadcb40abd4fb42946c30130ac5bd2983d74eb893e410b

Observation 5aeb4ef5-c030-4ff3-a977-6aa2a4c68dd7 · outbound

This paper cites DeepSeek 3FS.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference DeepSeek 3FS

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:46.087572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:45.106737Z digest=sha256:a314a26b1b0293d2a8e8c04bb3c4e9b321dd3a05fe8c42f35fd2d15bd8bb6a11

Observation 75256a89-6565-4e20-b516-4c4f2b4578ab · outbound

This paper cites IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:45.827026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:45.218533Z digest=sha256:4f05504e44aa9d42eea9520003ed0bf516323059722dd1b48ccc6a132c580bb4

Observation e12c1413-4397-4f13-abee-b623222209de · outbound

This paper cites Exploring cxl-based kv cache storage for llm serving,.

Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference Exploring cxl-based kv cache storage for llm serving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:45.549127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:45.314876Z digest=sha256:c3c12d6f08958a9a6acadd1aa0f552ee1c752d3026239d43620daa82c73fe520

Pith citing papers

No inbound Pith citation observations are available.