Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T09:50:23.266920Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 100 of 139 outbound references and 0 inbound Pith citation observations for arXiv:2607.02574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T09:50:23.266920Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 139 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cad7c44f-6a41-4c69-aaa3-a1985649e192 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving J., Soloveychik, I., and Kamath, P.Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems(2024)
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c656e6b-f60e-46e1-bd76-1f0a4060e388 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving S., Tumanov, A., and Ramjee, R.Taming throughput-latency tradeoff in llm inference with sarathi-serve
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c91eeed-246f-4938-89dd-1552cefc34ba · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f955fdc-3b0d-41d4-af45-40dc0961e38c · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8722d4cb-ccf0-4bce-a07b-6d2e7df9d5a6 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Y., Rajbhandari, S., Awan, A
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef42ae8-4920-4492-96e8-bebf2ea5d6be · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7089685c-341b-4297-b2cd-241b3c7b0ec3 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Longformer: The Long-Document Transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3b4bab-c379-4bc9-98ad-e1d195f3e73d · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Striped Attention: Faster Ring Attention for Causal Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d902cb1-6e55-4643-9e49-9c2ac65fad06 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d5b39d-f231-402e-9f12-58b019176a50 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d75ca09-a319-4dc8-ab7c-907e9d4fa4a4 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d09a5b7-3b94-4935-b244-bd4cddaa50f9 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c242c8-f385-4c08-b22e-1c3b79b89894 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a05479f-0a7f-434c-8ee2-b39f8233ba8e · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving KVDirect: Distributed Disaggregated LLM Inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97c3075-4db8-4c09-88ee-fae871c3836a · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Extending Context Window of Large Language Models via Positional Interpolation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42ea428-ea7c-497a-ad3c-3f35b4acc210 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InInternational Conference on Learning Representations(2024), vol
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2cc121-2ccd-4ac6-8b7d-45e21e2fef8e · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Generating Long Sequences with Sparse Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e4f6985-b4d9-47d7-82f9-cc1ccfff3fb4 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving [19]Corporation, A.Amd instinct mi300 series architecture.Technical whitepaper(2023)
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f3fa9ec-a9bd-4ee9-af45-e50297ab9e55 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InInternational Conference on Learning Representations(2024), vol
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85546b09-60a4-4c9e-800d-abe3fc17289d · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963a5b70-d258-4286-a67c-65e2761f1c9c · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1608e6ca-b4d6-4e66-970f-76d12497b53a · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 573f8ee7-c0a4-4b9b-97df-9027a41a8792 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving LongNet: Scaling Transformers to 1,000,000,000 Tokens
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbe4e20-5419-482e-b6d4-3fe02b480b56 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59e9c71-c4f4-41f5-a174-fe493555aa17 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b603132-cb31-4a85-b3d4-a402822a5efd · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e1b391-4fc2-4df7-9ec7-8042f205e8fb · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228d47dd-d854-4a54-bd31-57e2e0538e73 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312742af-27cc-42f4-9eea-172d1f55affc · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e50066-2dd8-45fa-a04e-412ce0b1869e · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving K.Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference.Advances in Neural Information Processing Systems(2024)
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43313fea-998d-4060-a4db-7e208c3744ab · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf437ec-0509-4136-888f-fc33deed1096 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd9a7c0-1059-4670-95b1-72526efb4668 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add754b7-c5b3-4ed8-ac62-14ab4bda340c · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Proceedings of Machine Learning and Systems(2022)
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ddacec-4761-4d33-ac38-863d092ac612 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving L., Khandelwal, A., and Zhong, L.Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems 6(2024), 325–338
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1142562-be7c-464e-927f-a92fbe08476d · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving The Llama 3 Herd of Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7698e042-ac99-4a38-93f8-a4fe2d6fe5ad · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c4e1f3-26aa-43aa-9720-461b72286256 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Efficiently Modeling Long Sequences with Structured State Spaces
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a4679e-1502-4402-a93e-8dbe20ee766e · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving G.Efficient memory disaggregation with infiniswap
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f05965a-be86-4cfe-8d50-ed665afaa379 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In International conference on machine learning(2020), PMLR, pp
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51a0686-9c62-4327-92f9-b133c37e797f · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6152210c-84fd-4605-a866-bb8536a07119 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving FastMoE: A Fast Mixture-of-Expert Training System
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e581dcb3-0272-4c4b-90fc-a1a5b587d221 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 531fff35-3744-4b4a-8c2a-31e98b63c230 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Training Compute-Optimal Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208c00e4-9b51-4ca6-a064-f638bf559b5f · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5843d529-260d-4b0d-9340-d3176d5aa4bf · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving W., Shao, Y
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd7edb7a-853e-4df5-93bf-7fe941c9f6ca · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ce9caf-d125-4573-859e-4df4ba102f0a · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357e8dc0-9690-4498-a1b2-ad0498d78dff · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087530ff-cec7-4866-8104-0a7a488b5d9e · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcda397-0e71-4801-bc54-f44097a94024 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Mixtral of Experts
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4aa7403-4525-49aa-8078-8c3835e9f9ee · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08698169-f112-45b2-a260-1d555cd15ccd · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving P/D-Serve: Serving Disaggregated Large Language Model at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd44819b-a4ab-40e8-b6d8-e641d436a082 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8e962e-f30c-442b-9bf0-228e2aebb6a8 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Hydragen: High-Throughput LLM Inference with Shared Prefixes
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdd29e4-e91d-495a-807c-9983e269bd1a · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dee23f8-e923-4207-aa37-1c33a573b168 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 2020 conference on empirical methods in natural language processing (EMNLP)(2020), pp
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea433842-8f30-462a-9a04-71edc885daf3 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval(2020), pp
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf85f5b-64cf-48d2-9976-b7fd0cb2b1a5 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Reformer: The Efficient Transformer
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d31af90-403a-4dd2-bada-ee8f025bb7b0 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving H., Gonzalez, J., Zhang, H., and Stoica, I.Efficient memory management for large language model serving with pagedattention
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6d477a-9f9b-4c36-8a99-5bae46cb18de · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 57th annual meeting of the association for computational linguistics(2019), pp
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211e8de5-93dc-4361-955d-4bb84908d578 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dce03ab-1acb-43aa-84fa-2735851764fd · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving T., and Rabusseau, G.Kq-svd: Compressing the kv cache with provable guarantees on attention fidelity.arXiv preprint arXiv:2512.05916(2025)
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc23d8a7-f710-4246-a953-66a7669715cd · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In International Conference on Machine Learning(2023), PMLR, pp
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ca30aa-0a60-4a97-8da3-41aaa7706bbc · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a248df5-aef7-489e-a79a-2dca67a3f63f · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving S., Hsu, L., Ernst, D., Zardoshti, P., Novakovic, S., Shah, M., Rajadnya, S., Lee, S., Agarwal, I., et al.Pond: Cxl-based memory pooling systems for cloud platforms
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef701a40-15eb-42eb-a2a4-373a290017d4 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65508025-0de0-4c58-be13-85faaa81ee79 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad158e2-5212-4bdd-ae10-a5ffbd292fbf · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving E., et al
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1beb5909-efba-4ad6-bee9-51a4df9b23de · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118eae73-320f-4f63-b75d-5cb40ac5aab7 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0421319-d752-4fe9-9cba-5113b796ec04 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ba35715-bbc0-49ee-869a-aeeb2aba232c · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 2026 ACM/SIGDA International Symposium on Field Programmable Gate Arrays(2026), pp
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ef4c7bd-23e1-402e-be26-1b56f38db33b · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf9a85e-7be7-4530-abc5-e0f4635963c1 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In International Conference on Learning Representations(2024), vol
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2a5b33-c4b9-4a95-ae6b-adc48afcef63 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving 32 Jie Li, Tongyang Wang, and Yong Chen
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06378a37-5429-4be8-a8f3-b6c1de6cfa53 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 714f0613-9669-4e7d-929c-14b59d3d71ea · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 4th International Conference on Artificial Intelligence and Intelligent Information Processing(2025), pp
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca62b75c-37a0-4718-b741-346f2264b383 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the ACM SIGCOMM 2024 Conference(2024), pp
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b22ac4-918e-410d-bd3a-069bc9f6b8df · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 462fbf52-a1df-431b-b5a0-4b74668bb214 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d5aabc-7126-4165-8efb-0d2a84b55720 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4db615-375d-4355-98a7-c6d1d72c3b4d · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving A., and Y ashunin, D
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a952099d-d4ab-4e7c-8d16-7dbead6f1312 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e040c55-e160-4981-9490-b0c2cb6606aa · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dfef44-92e1-4fd9-b9f5-92dd9c6fd3ea · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Landmark Attention: Random-Access Infinite Context Length for Transformers
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1c1b3ba-d766-461b-8d3e-e79258171d03 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the international conference for high performance computing, networking, storage and analysis (2021), pp
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b714451a-be66-4a3f-b7c6-4554e09617e3 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2c30df-c0c6-4002-ab4c-4949fc289c52 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0acc352b-3cec-4e36-a356-6ba23a0d4139 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving M., Stratmann, E., and Stutsman, R.The case for ramclouds: Scalable high-performance storage entirely in dram.ACM SIGOPS Operating Systems Review 43, 4 (2010), 92–105
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3edabbc9-277b-49cf-91b6-d8325537e2c5 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA)(2024), IEEE, pp
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eedc6ea-86c3-4eb4-9621-d9512f083101 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InInternational Conference on Learning Representations(2024), vol
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e6404b2-4bd0-4c60-8010-1559a9d0e226 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Efficiently scaling transformer inference.Proceedings of machine learning and systems 5(2023), 606–624
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132596d3-6ab5-43cd-822d-b83f3f8ecc00 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(2025), pp
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e68687d-7b86-4dab-aafd-3ac83fddfb57 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving DRackSim: Simulator for Rack-scale Memory Disaggregation
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 480e6ef2-f300-43c2-aeb1-f9e7c162968f · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff09ccd4-b4d6-4775-a6fb-60b9c473a77c · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Y., Awan, A
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f7cc40-a63f-43af-a802-1a258a055db3 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving InSC20: international conference for high performance computing, networking, storage and analysis(2020), IEEE, pp
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25231359-91c6-47ae-ac93-4234e7f1fc66 · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Unresolved cited work
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42a42c12-72e9-43f7-a0bb-84c603b381ec · outbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.