Pith. sign in

Paper Citation Record · LEDGER

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving

As of 4 August 2026, this Paper Citation Record lists 100 of 139 outbound references and 0 inbound Pith citation observations for arXiv:2607.02574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02574 v1

Coverage vector

measured 100 of 139 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:50:23.266920Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 139 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cad7c44f-6a41-4c69-aaa3-a1985649e192 · outbound

This paper cites J., Soloveychik, I., and Kamath, P.Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems(2024).

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving J., Soloveychik, I., and Kamath, P.Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems(2024)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:d511b78e28f817cbcfb80e2cc48835871d4926cef60302bed57cb8d74179a2b5

Observation 2c656e6b-f60e-46e1-bd76-1f0a4060e388 · outbound

This paper cites S., Tumanov, A., and Ramjee, R.Taming throughput-latency tradeoff in llm inference with sarathi-serve.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving S., Tumanov, A., and Ramjee, R.Taming throughput-latency tradeoff in llm inference with sarathi-serve

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ca1d54a7ab0fff5b3ee88f5a1b3896ac6e2298577661ae0c38773028ca98c559

Observation 4c91eeed-246f-4938-89dd-1552cefc34ba · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:def9f7c36fa4d972b64d64d570a03638c527fc19af144d3d8ee7b063ff97385e

Observation 7f955fdc-3b0d-41d4-af45-40dc0961e38c · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:bd331c049253f16c5f8c168bca6269ff94cb5836662bd8258ea281d09f768fb1

Observation 8722d4cb-ccf0-4bce-a07b-6d2e7df9d5a6 · outbound

This paper cites Y., Rajbhandari, S., Awan, A.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Y., Rajbhandari, S., Awan, A

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:672259b0ec8a50d767f77e7b9f355148aec1ce495f56f3d3273fb11b39625a12

Observation 3ef42ae8-4920-4492-96e8-bebf2ea5d6be · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:7c3f5c2d5463f3ec7a9fee40c76dfa02562bcf6e7a1fdff72eeaba068bd72a8e

Observation 7089685c-341b-4297-b2cd-241b3c7b0ec3 · outbound

This paper cites Longformer: The Long-Document Transformer.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Longformer: The Long-Document Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:d432ea8634dc099e01c2a79071c524939b4594b955c73e05c8d1704c8f535b3d

Observation 1e3b4bab-c379-4bc9-98ad-e1d195f3e73d · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Striped Attention: Faster Ring Attention for Causal Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6120d5a8ce5de5347a58556a85304e1dba21c1b23cc1f9615a0cb61de3df3723

Observation 2d902cb1-6e55-4643-9e49-9c2ac65fad06 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:84c7de3690acb96ae825a77421fc1b42f83fd9529a2d28e475a72177874fda15

Observation 28d5b39d-f231-402e-9f12-58b019176a50 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:de5320097d0963020347eae167e7e242f08402073a41ebe77b9b18e70ea844d9

Observation 6d75ca09-a319-4dc8-ab7c-907e9d4fa4a4 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:c0d523c72bf39dcb7cf3b57dcbd1f52aecfe46ae03a394b22f681d8ff1831b52

Observation 7d09a5b7-3b94-4935-b244-bd4cddaa50f9 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:5b31c310969a9d103134439e57bedab67ae88953c84ee32c7462fdd28fb823f2

Observation 27c242c8-f385-4c08-b22e-1c3b79b89894 · outbound

This paper cites Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:8f5ccee02cc1bb0480b0ef8d8e1aa1596e31c86a60a2913a64d177a027e45128

Observation 4a05479f-0a7f-434c-8ee2-b39f8233ba8e · outbound

This paper cites KVDirect: Distributed Disaggregated LLM Inference.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving KVDirect: Distributed Disaggregated LLM Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:f48ccd2393dd7c17c3f8ee2db09e8ead150d54fca26c6f4549f8e373276d75df

Observation e97c3075-4db8-4c09-88ee-fae871c3836a · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Extending Context Window of Large Language Models via Positional Interpolation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4997450ba0333ebd09b7092f1400c0b94786c273332a44b2053d0dfc61383fe8

Observation d42ea428-ea7c-497a-ad3c-3f35b4acc210 · outbound

This paper cites InInternational Conference on Learning Representations(2024), vol.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InInternational Conference on Learning Representations(2024), vol

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:2bf0a8b8ae02b35d9a1ceacae175277f55ea4c27110d0940288db132ee5ff62a

Observation 0b2cc121-2ccd-4ac6-8b7d-45e21e2fef8e · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Generating Long Sequences with Sparse Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:1cd2c93bb3d32af452feb7bba129d4fc221449ceb3d12a5218e96cbb63abe440

Observation 0e4f6985-b4d9-47d7-82f9-cc1ccfff3fb4 · outbound

This paper cites [19]Corporation, A.Amd instinct mi300 series architecture.Technical whitepaper(2023).

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving [19]Corporation, A.Amd instinct mi300 series architecture.Technical whitepaper(2023)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:3866902212d2b15a1213e738727f9c5d70bf27330793cae10104cae555297323

Observation 7f3fa9ec-a9bd-4ee9-af45-e50297ab9e55 · outbound

This paper cites InInternational Conference on Learning Representations(2024), vol.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InInternational Conference on Learning Representations(2024), vol

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:53d23ae09ba011b88ba7a07622eb381c6c8b5337f342746fe979e7937c3b2709

Observation 85546b09-60a4-4c9e-800d-abe3fc17289d · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:563c9f34dd3607857b2b6df0d69bf993fbe13386299abb337f6dca5d5ead3364

Observation 963a5b70-d258-4286-a67c-65e2761f1c9c · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:e598bb3fb7f3b6a0954e80113cb2fa4d40fa06b3ce2d828090eeb0b49716e88c

Observation 1608e6ca-b4d6-4e66-970f-76d12497b53a · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ce95d6f0ebe34cf454f4e28af4f7b83e56be5e8200af10df84b4a4c92567a5f9

Observation 573f8ee7-c0a4-4b9b-97df-9027a41a8792 · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:c1334f2d74b207b93c9da438c0774315927babb4cbd045f4019cc8e81671a93c

Observation 9dbe4e20-5419-482e-b6d4-3fe02b480b56 · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:bb612c0a3516e89dfb8976afc2e218532f868c7e844b565c3b4a19226393e8c2

Observation b59e9c71-c4f4-41f5-a174-fe493555aa17 · outbound

This paper cites Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:f2ae6c269364c5d9b8ceaa6d5e2afc90bb691492c84583096e529764a27eb9f0

Observation 8b603132-cb31-4a85-b3d4-a402822a5efd · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:5df1c06ffaee733fe0135b39a85a50ef9bfaf4bb6ba77f74879ba6c025e66298

Observation 82e1b391-4fc2-4df7-9ec7-8042f205e8fb · outbound

This paper cites M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:0f503b2cc0eb38102019929b55323aef34995b9238d49e2386d43b1df00e026d

Observation 228d47dd-d854-4a54-bd31-57e2e0538e73 · outbound

This paper cites MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:76b9dec46226ec02efb3a0ee57804ffb41f2e1ceb54d171122496a2d4a88b4bb

Observation 312742af-27cc-42f4-9eea-172d1f55affc · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:3eb5015fa06f7cef775a62e0a0d39f7f1112e5658e65f0a3f2045df9e0405a5d

Observation 59e50066-2dd8-45fa-a04e-412ce0b1869e · outbound

This paper cites K.Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference.Advances in Neural Information Processing Systems(2024).

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving K.Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference.Advances in Neural Information Processing Systems(2024)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:14ab9e4fa219cebebc9af101e0946e3e1f3489ee122f0d09812412f9ff13babe

Observation 43313fea-998d-4060-a4db-7e208c3744ab · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:52389df000004eb4ec7e36a2062610004be861e5731843e51010e3457c43bb30

Observation 4cf437ec-0509-4136-888f-fc33deed1096 · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:63740a91b073f3cd1751515669a94aac685b79177a4632f46fd7704d477f74db

Observation 9fd9a7c0-1059-4670-95b1-72526efb4668 · outbound

This paper cites Break the Sequential Dependency of LLM Inference Using Lookahead Decoding.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:0759a598dd1fb1b555ab0a5abad4ba2f5848f2a6fd25cef83168977694e6f61f

Observation add754b7-c5b3-4ed8-ac62-14ab4bda340c · outbound

This paper cites Proceedings of Machine Learning and Systems(2022).

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Proceedings of Machine Learning and Systems(2022)

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:5332fa78b207bb7b2d4fea31b3dd04bd8f134ab3aa38ef08dbc1ede99f42878e

Observation a0ddacec-4761-4d33-ac38-863d092ac612 · outbound

This paper cites L., Khandelwal, A., and Zhong, L.Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems 6(2024), 325–338.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving L., Khandelwal, A., and Zhong, L.Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems 6(2024), 325–338

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:a48a03621b40ddfa3300ac6aa45b098c0cf81167c3a3bc2197eed3f49e71a8c2

Observation b1142562-be7c-464e-927f-a92fbe08476d · outbound

This paper cites The Llama 3 Herd of Models.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving The Llama 3 Herd of Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:0ef2501459065d4bf2341c03813374ef61c49bb81dba59e0113a240c0a8afaa1

Observation 7698e042-ac99-4a38-93f8-a4fe2d6fe5ad · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:93b7707182e71cbb1bf3d33e14ffbf088201ced4f583854085d229d144daaaca

Observation 29c4e1f3-26aa-43aa-9720-461b72286256 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Efficiently Modeling Long Sequences with Structured State Spaces

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:2c77f514708ebea247edb0c4b7155441b9dcb84140bed63b48d53a4cb8af984d

Observation c6a4679e-1502-4402-a93e-8dbe20ee766e · outbound

This paper cites G.Efficient memory disaggregation with infiniswap.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving G.Efficient memory disaggregation with infiniswap

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4079000bd188a11898490c41748b7faf417ba88d919987f47f282394513788f5

Observation 2f05965a-be86-4cfe-8d50-ed665afaa379 · outbound

This paper cites In International conference on machine learning(2020), PMLR, pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In International conference on machine learning(2020), PMLR, pp

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:99e0fa4454f96e9b179f22db2a351477739d5252331f58cdf5d34cf90e9daec2

Observation f51a0686-9c62-4327-92f9-b133c37e797f · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6f95d771701d8fd26206ed55c9302ac0d3b140bafd08fb11e24d96724f44fc6c

Observation 6152210c-84fd-4605-a866-bb8536a07119 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving FastMoE: A Fast Mixture-of-Expert Training System

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:c6e39c3483efa5d13bc588007cbb05dec7348c9f04bd9bbde09a8ab00a2668c3

Observation e581dcb3-0272-4c4b-90fc-a1a5b587d221 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:a9945d90c84d9cbf773bb4ca239c9287e4d8830bc256c4ff3b297d31612d5ab6

Observation 531fff35-3744-4b4a-8c2a-31e98b63c230 · outbound

This paper cites Training Compute-Optimal Large Language Models.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Training Compute-Optimal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:5d8ed4abb2aa07cb048d42534c52a425b33cca0350e669fa9ca4ce574ab3052c

Observation 208c00e4-9b51-4ca6-a064-f638bf559b5f · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ec65faa90827c3383285e0f0b22120d8e66869f9ee3810fac6abc8e01d169649

Observation 5843d529-260d-4b0d-9340-d3176d5aa4bf · outbound

This paper cites W., Shao, Y.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving W., Shao, Y

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:385108efc843b516539481131c86d6b15682b9a4909c32103d9733f5205f45d6

Observation fd7edb7a-853e-4df5-93bf-7fe941c9f6ca · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4c7725a895b196eb0d55484dc4fbe7a2db978113b7190354a6d4fad460bc2c46

Observation a8ce9caf-d125-4573-859e-4df4ba102f0a · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:978ec22b00df4933ffe755047c72233cc2df9cac76d73d8cc47e2b087db3dd3c

Observation 357e8dc0-9690-4498-a1b2-ad0498d78dff · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:02bf7019bd1ae1d13f3d6df9643992532f2949c965e3e947dfe7cd86d836e8bf

Observation 087530ff-cec7-4866-8104-0a7a488b5d9e · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:239226c5dd220eedaa4865db159dad897405524607d76cacafec6def6cb73bbe

Observation fdcda397-0e71-4801-bc54-f44097a94024 · outbound

This paper cites Mixtral of Experts.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Mixtral of Experts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:60496734a1ee28dc64daeec6ebdc527464b0e9c69ccb57ab7a3dbbd31c71d0fa

Observation e4aa7403-4525-49aa-8078-8c3835e9f9ee · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:c00d0d9b0e93eee8d42c3a9f60fc3d085bf97f5e7162ab5e22f18ae7ffd65a1f

Observation 08698169-f112-45b2-a260-1d555cd15ccd · outbound

This paper cites P/D-Serve: Serving Disaggregated Large Language Model at Scale.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:249149ed181a9ad01660a13456c41d6b9c2cecf25cc5b5b073dd2a5644c453cd

Observation dd44819b-a4ab-40e8-b6d8-e641d436a082 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ee71bbefa35b361974a9a104cdc0f6aeeffb8c27335464ec3ab3a0e86d913dc8

Observation ec8e962e-f30c-442b-9bf0-228e2aebb6a8 · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:5cb897c4da2cff156f0e3e170d3461bf65a5dd086febd6f49f5aafab7244bb03

Observation 7cdd29e4-e91d-495a-807c-9983e269bd1a · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:e31c4b8cfae3ae734ab19237ecf3fbced1b7fc97a558dd04007f8b6a1e9b0819

Observation 4dee23f8-e923-4207-aa37-1c33a573b168 · outbound

This paper cites InProceedings of the 2020 conference on empirical methods in natural language processing (EMNLP)(2020), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 2020 conference on empirical methods in natural language processing (EMNLP)(2020), pp

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:0764c8abaecab8aebb1cd381f8276cfdcb525d0a0be69fa06ff07a208dc736d7

Observation ea433842-8f30-462a-9a04-71edc885daf3 · outbound

This paper cites InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval(2020), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval(2020), pp

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6386eca4b8e0453501409fba7a560342522941b7df27922d0a80894dd48443b7

Observation 6cf85f5b-64cf-48d2-9976-b7fd0cb2b1a5 · outbound

This paper cites Reformer: The Efficient Transformer.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Reformer: The Efficient Transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:0105a5e9df7bac6e5a2cbcec70f6ccce4b38b30f1de15528346ba4b4d863d5b7

Observation 7d31af90-403a-4dd2-bada-ee8f025bb7b0 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.Efficient memory management for large language model serving with pagedattention.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving H., Gonzalez, J., Zhang, H., and Stoica, I.Efficient memory management for large language model serving with pagedattention

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:57bf148923ecebaa31f680770e327e068eff1f8f55ff81d55be44f5233e68038

Observation ed6d477a-9f9b-4c36-8a99-5bae46cb18de · outbound

This paper cites InProceedings of the 57th annual meeting of the association for computational linguistics(2019), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 57th annual meeting of the association for computational linguistics(2019), pp

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:cf5c14dc867e85b162f2c272c3a32bee0a86b05b7455ee5114ecfee7b1d75304

Observation 211e8de5-93dc-4361-955d-4bb84908d578 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:dcf0ddc89df0baf1cf1406e4c3fd7d49e63e07bd74653f1c398e6eb3e7f2882e

Observation 2dce03ab-1acb-43aa-84fa-2735851764fd · outbound

This paper cites T., and Rabusseau, G.Kq-svd: Compressing the kv cache with provable guarantees on attention fidelity.arXiv preprint arXiv:2512.05916(2025).

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving T., and Rabusseau, G.Kq-svd: Compressing the kv cache with provable guarantees on attention fidelity.arXiv preprint arXiv:2512.05916(2025)

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:89b5450104220601276c7b2767c444ec42df2212d9bef2509fa4d8e58dedc65c

Observation dc23d8a7-f710-4246-a953-66a7669715cd · outbound

This paper cites In International Conference on Machine Learning(2023), PMLR, pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In International Conference on Machine Learning(2023), PMLR, pp

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:26b8a2b9f4f31d942d5c64efb3066570fd6daef385bdc4ac3bd491b5abf7fef8

Observation a4ca30aa-0a60-4a97-8da3-41aaa7706bbc · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:b12954c0fbde142613fe2ee0f3b69476aeeb18657af8b4b5e60c443846c98c09

Observation 2a248df5-aef7-489e-a79a-2dca67a3f63f · outbound

This paper cites S., Hsu, L., Ernst, D., Zardoshti, P., Novakovic, S., Shah, M., Rajadnya, S., Lee, S., Agarwal, I., et al.Pond: Cxl-based memory pooling systems for cloud platforms.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving S., Hsu, L., Ernst, D., Zardoshti, P., Novakovic, S., Shah, M., Rajadnya, S., Lee, S., Agarwal, I., et al.Pond: Cxl-based memory pooling systems for cloud platforms

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:e444d3aa4a58687abd1f409c63652a00e9552ebc8f504623e412555b686bb87b

Observation ef701a40-15eb-42eb-a2a4-373a290017d4 · outbound

This paper cites A Survey on Large Language Model Acceleration based on KV Cache Management.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving A Survey on Large Language Model Acceleration based on KV Cache Management

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:c6a344adcdd206217a21aa095d5958df052b98c2aa3c857428d83032d5dd5c6d

Observation 65508025-0de0-4c58-be13-85faaa81ee79 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:baf9161cca52c5b08ee5f21ad158aaa562468f1b7d5f36fe67d5764ec66f3f12

Observation aad158e2-5212-4bdd-ae10-a5ffbd292fbf · outbound

This paper cites E., et al.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving E., et al

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:77275ebad935a626ae0d29e0452773e8bb66ff10778f1d37be84ad7d9cd71a84

Observation 1beb5909-efba-4ad6-bee9-51a4df9b23de · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:fa0927cbaba3186174837b480b37d98b51415c8fa4d080fc30322df0fb741dd9

Observation 118eae73-320f-4f63-b75d-5cb40ac5aab7 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:c196426df2acf396ba4dbc4ce555bb5b0171aa4d435e9ec92dcafa83d0c40d4d

Observation b0421319-d752-4fe9-9cba-5113b796ec04 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:2bc9ac8dd5a9b37376cbd9c69eb2f209115a636262094f597c71d4e505143ebe

Observation 4ba35715-bbc0-49ee-869a-aeeb2aba232c · outbound

This paper cites InProceedings of the 2026 ACM/SIGDA International Symposium on Field Programmable Gate Arrays(2026), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 2026 ACM/SIGDA International Symposium on Field Programmable Gate Arrays(2026), pp

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:3329d044515b9002ba7dce880410281004c312f4af4bc424c289d97013a507ee

Observation 6ef4c7bd-23e1-402e-be26-1b56f38db33b · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:abd7f758e562b32f35fa9401efc648f2d98030b2e94792678e1db8c156df0ef6

Observation dbf9a85e-7be7-4530-abc5-e0f4635963c1 · outbound

This paper cites In International Conference on Learning Representations(2024), vol.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In International Conference on Learning Representations(2024), vol

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:b4b0c272d80e036bbf12234695ce5ed7f97b43ce7fa333e61dd11f2e8dd3d4fa

Observation ce2a5b33-c4b9-4a95-ae6b-adc48afcef63 · outbound

This paper cites 32 Jie Li, Tongyang Wang, and Yong Chen.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving 32 Jie Li, Tongyang Wang, and Yong Chen

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:131a19576e67f1ce676fe3143e5b206aa9c2137116596767ae54306d68330236

Observation 06378a37-5429-4be8-a8f3-b6c1de6cfa53 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:5ffd3b83757d32e45160db62038d4814ff03f62381583170929b434bb7d929f0

Observation 714f0613-9669-4e7d-929c-14b59d3d71ea · outbound

This paper cites InProceedings of the 4th International Conference on Artificial Intelligence and Intelligent Information Processing(2025), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 4th International Conference on Artificial Intelligence and Intelligent Information Processing(2025), pp

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:de1f282ad311dcaddd7c1504ea494e0399c5d61e9839f274daec9100ffacbc5e

Observation ca62b75c-37a0-4718-b741-346f2264b383 · outbound

This paper cites InProceedings of the ACM SIGCOMM 2024 Conference(2024), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the ACM SIGCOMM 2024 Conference(2024), pp

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:e4660673ef9cdb89d964a1c00be5a9eb288e56f82df108e99972c2e37cff7afd

Observation e5b22ac4-918e-410d-bd3a-069bc9f6b8df · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:ca453a17640b1bd4df8ae28eff9ba5bcbf718869cd6f1d92fa5527020fe0b2b3

Observation 462fbf52-a1df-431b-b5a0-4b74668bb214 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:703947da9f6eb2d3e0640c28620aaee62f63f94e8aa8698474c6dcc8ef0345fc

Observation b1d5aabc-7126-4165-8efb-0d2a84b55720 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:fefe19dcc474fe56b92384e2217f0d89bb184ecfea08d75ec1050a8b9dd9bebb

Observation 7b4db615-375d-4355-98a7-c6d1d72c3b4d · outbound

This paper cites A., and Y ashunin, D.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving A., and Y ashunin, D

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:1e97557c789d88db2427a9a251e6e943d98e29314e88bcf1d56d88ad86b54298

Observation a952099d-d4ab-4e7c-8d16-7dbead6f1312 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:741c5c195f430744471f30ba65b7c33d6b8862f0d4e36013633856c6004051d2

Observation 3e040c55-e160-4981-9490-b0c2cb6606aa · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:6dbbda8f39aa3edb39ff967a675417fc3f75d5f991efe703253ad12f15b4d7c2

Observation 31dfef44-92e1-4fd9-b9f5-92dd9c6fd3ea · outbound

This paper cites Landmark Attention: Random-Access Infinite Context Length for Transformers.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Landmark Attention: Random-Access Infinite Context Length for Transformers

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4272ceb5d8c629df0234a5fea8a6a295f21197adb5cc26e1f744ebe04f17a267

Observation c1c1b3ba-d766-461b-8d3e-e79258171d03 · outbound

This paper cites InProceedings of the international conference for high performance computing, networking, storage and analysis (2021), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the international conference for high performance computing, networking, storage and analysis (2021), pp

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:95970997c09fbcaf5acfea5debede04945b60c2755cec0f9ddd563b84912c81d

Observation b714451a-be66-4a3f-b7c6-4554e09617e3 · outbound

This paper cites Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:37aedbe7e8a1cb5ffee0d2031ab0a0f1916967263c6e1eb0fd29fef07f0c837b

Observation 7e2c30df-c0c6-4002-ab4c-4949fc289c52 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:8d31a87d67812de944a488a7738e484336e4af85109b3dda4fed7c0038157765

Observation 0acc352b-3cec-4e36-a356-6ba23a0d4139 · outbound

This paper cites M., Stratmann, E., and Stutsman, R.The case for ramclouds: Scalable high-performance storage entirely in dram.ACM SIGOPS Operating Systems Review 43, 4 (2010), 92–105.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving M., Stratmann, E., and Stutsman, R.The case for ramclouds: Scalable high-performance storage entirely in dram.ACM SIGOPS Operating Systems Review 43, 4 (2010), 92–105

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:25aae1bc0f859fdda433e4353bd80a4e088ad0416626f84312d7e06eba9956a4

Observation 3edabbc9-277b-49cf-91b6-d8325537e2c5 · outbound

This paper cites In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA)(2024), IEEE, pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA)(2024), IEEE, pp

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:e4d2536049a7d0d035be7dcbf073eb21dca97100c1e7e670baf49c4e1926c59a

Observation 7eedc6ea-86c3-4eb4-9621-d9512f083101 · outbound

This paper cites InInternational Conference on Learning Representations(2024), vol.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InInternational Conference on Learning Representations(2024), vol

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:f1c54819a7f6bce52bf440bdaa890a73ef88066e426b724bd0d1accafcfe4073

Observation 9e6404b2-4bd0-4c60-8010-1559a9d0e226 · outbound

This paper cites Efficiently scaling transformer inference.Proceedings of machine learning and systems 5(2023), 606–624.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Efficiently scaling transformer inference.Proceedings of machine learning and systems 5(2023), 606–624

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:7bcf9926f1f2dacd03f7fd22a53bc365618defacf0ab790bf42bb334d3546101

Observation 132596d3-6ab5-43cd-822d-b83f3f8ecc00 · outbound

This paper cites InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(2025), pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(2025), pp

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:3d0080c1bc74fcfefa4c8f1509fc5250aa87fb6a41329bbf81b4ce825617d3a6

Observation 8e68687d-7b86-4dab-aafd-3ac83fddfb57 · outbound

This paper cites DRackSim: Simulator for Rack-scale Memory Disaggregation.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DRackSim: Simulator for Rack-scale Memory Disaggregation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:314d8c95c7dde3a6663035ee9d210c9154d959a02bd57a7f91c94559cdd4a6a6

Observation 480e6ef2-f300-43c2-aeb1-f9e7c162968f · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4b51d4b25b1d475c999ea0be57ad8c7968f7ae47ed918f09ca133b5610d43795

Observation ff09ccd4-b4d6-4775-a6fb-60b9c473a77c · outbound

This paper cites Y., Awan, A.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Y., Awan, A

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:d62b3b5f41e6aa1e472f22616613e3e9b3ca81d984940d6af724c77b00f49b97

Observation 81f7cc40-a63f-43af-a802-1a258a055db3 · outbound

This paper cites InSC20: international conference for high performance computing, networking, storage and analysis(2020), IEEE, pp.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InSC20: international conference for high performance computing, networking, storage and analysis(2020), IEEE, pp

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:a7af9437707fe4ebebbfd348de483e3a386fb8ffd392270a56bd8077aded610e

Observation 25231359-91c6-47ae-ac93-4234e7f1fc66 · outbound

This paper cites an unresolved cited work.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:9a6cf43420937c91c1f6f83436952d882da7c66de4096557d6fff5af00c2c5e2

Observation 42a42c12-72e9-43f7-a0bb-84c603b381ec · outbound

This paper cites GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:537fef7070d5b909da0c128126dcfd1d467ee7995b3be928207d5209780963a7

Pith citing papers

No inbound Pith citation observations are available.