REVIEW 3 major objections 5 minor 76 references
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read KV Cache reuse makes LLM inference an internet-scale content-management problem; the paper proposes decoupling compute and cache storage across clouds.
desk verdict Timely vision paper for cross-cloud KV-cache distribution, worth reading for the framing and the survey; the cost model has a fixable unit bug and the reuse premise is untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the KV Cache treated as mobile inference state, together with the break-even cost identity in Eq. (2). That identity separates the model-side ratio $\mathit{PrefillTFLOPs}/\mathit{KVMB}$ from the infrastructure-and-application-side ratio $(\mathit{DeviceCost}_{\mathrm{hr}}\cdot K + \mathit{TransferCost})/(\mathit{GPUCost}_{\mathrm{hr}}/\mathit{PeakFlopRate})$. When the infrastructure side is smaller than the model side, storing and transferring the cache beats recomputing; otherwise recompute wins. This converts a storage-versus-compute decision into a network-aware, content-distribution decision and grounds the paper's proposed global control plane for lookup, place
What would settle it
A measurement study of production LLM and agent traces that computes the fraction of KV Cache reuse that is exact-prefix versus partial or semantic, the fraction that crosses cloud or region boundaries, and the distribution of reuse horizons. If cross-boundary reusable traffic is negligible, or horizons are consistently shorter than the time and cost of transfer, the paper's cost identity will almost never select internet transfer over recompute, falsifying its central economic premise.
Extended reading notes
Core claim
The paper's central claim is that the KV Cache—the tensor state storing the keys and values of a prompt's attention computation—is not merely a serving optimization artifact but a compact, persistent, steerable representation of contextual knowledge that can be shared, moved, and reused across requests, applications, and even models. Because inference workloads now span agents, retrieval, tool use, and multi-turn sessions, overlapping contexts recur naturally; the paper argues that this makes the location of compute relative to cached state a network-level decision rather than a datacenter-local one. The claim is made concrete by a cost identity: storing and transferring a reusable KV Cache
Load-bearing premise
The paper's internet-scale economics rest on the unmeasured premise that real LLM and agent workloads frequently reuse overlapping context across requests, tenants, and regions, with reuse horizons of seconds to hours; if reuse is mostly exact-prefix and confined to a single datacenter, the cost identity will not favor transferring KV Caches over the Internet.
Editorial extensions
If this is right
- A global KV Cache control plane—lookup, placement planning, prefetching, and recompute fallback—becomes necessary infrastructure for large-scale LLM inference.
- Network bandwidth and transfer price become first-class inference cost inputs, on par with GPU cost and storage cost, when deciding whether to store, fetch, partially reconstruct, or recompute a cache.
- Cache management should optimize for reducing miss-rate rather than hit-rate, since a miss on a long context is disproportionately expensive; cache blending and chunk-level reuse extend reuse beyond exact-prefix matches.
- KV Cache placement becomes adaptive across storage tiers and providers, with reuse horizon determining hot, warm, or cold placement and with graceful degradation to recomputation when transfer fails or becomes too expensive.
- Cross-provider agreements on pricing, ownership, privacy, and failure handling, plus standardized protocols, are required before cached inference state can move freely across administrative boundaries.
Reading between the lines
- A production-trace measurement of reuse frequency $K$ across requests, tenants, and regions would directly test the vision: if cross-boundary reuse is rare, the cost identity will almost never favor internet-scale transfer, and the paper's proposal collapses into today's intra-datacenter prefix caching.
- The same cost identity, applied at smaller scale, also justifies network-aware KV Cache placement within a single datacenter, so the paper's decoupling logic has value even if the full cross-cloud vision does not materialize.
- Model-side KV Cache compression changes the economics non-monotonically: smaller caches make transfer cheaper, but they also make recompute cheaper per token, so the optimal topology depends on the relative rates of compression and prefill-cost reduction over time.
- If KV Cache portability across providers becomes standardized, inference cost could become commoditized similarly to CDN bandwidth, raising new questions about cache ownership, privacy of cached prompts, and regulatory treatment of inference-state transit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a vision/position paper. It argues that LLM inference's KV cache has become an internet-scale content-management problem rather than a local compute-storage tradeoff, because contexts can be reused across requests, agents, and regions. The authors propose an "Internet for the KV Cache" with decoupled compute and storage across clouds and datacenters, a global control plane (KV Cache Directory, Lookup Service, Placement Planner, Prefetch Scheduler), and placement/reuse decisions driven by model-side, infrastructure-side, and application-side metrics. The argument is supported by a cost model in Eqs. (1)-(2), a storage-versus-recompute decision table (Table 1), a motivating numerical example (1 GB KV cache transfer vs. recompute), and a broad survey of model-side, compression, and system-side advances. The paper concludes with challenges in network-aware cache movement, adaptive control-plane design, and cross-provider coordination.
Significance. If the central premise were established, this would be a timely and valuable reframing: it connects KV cache management to content distribution, makes network bandwidth and transfer price first-class inference resources, and proposes concrete control-plane abstractions. The literature survey and the explicit cost model are useful starting points, and the paper explicitly names falsifiable decision metrics. However, the paper provides no measured reuse distributions, and the quantitative decision rule in Eq. (2) has a unit inconsistency in the definition of K. The central empirical premise—that cross-boundary, semantically overlapping context reuse is frequent enough to amortize wide-area transfer and storage—is asserted rather than demonstrated. The paper should be revised to either provide evidence from real traces or clearly mark the quantitative claims as conditional on unverified reuse assumptions.
major comments (3)
- [§3.2, Eqs. (1)-(2), Table 1] K is defined as "frequency of KV Cache reuse," but Table 1 assigns K the values "1 day" and "1 month," which are durations, not frequencies. With DeviceCost_hr in $/hr/MB, the term DeviceCost_hr * K * KV_MB has units of $/hr if K is dimensionless, making the ratio in Eq. (1) dimensionally inconsistent (1/hr). If K is instead intended as a retention interval in hours, the notation and interpretation must be changed, and the storage cost should be amortized over the number of reuses during that interval. As written, the break-even condition and all Table 1 decisions are ill-defined. This is load-bearing because Table 1 is the paper's main quantitative demonstration that "multiple correct solutions" exist. Please correct the definitions and recompute Table 1.
- [§1, §3.3, Table 1] The paper's motivating premise is that real LLM and agent workloads exhibit frequent, semantically reusable context overlap across requests, tenants, and regions, with reuse horizons of seconds to hours. No measurements or cited traces support this. Section 3.3 itself concedes that prefix caching is brittle to small input changes, so the argument depends on non-prefix semantic reuse (e.g., CacheBlend-style chunk fusion or agent-memory compaction) working at scale across administrative boundaries; no evidence is given. Table 1 presupposes the conclusion by setting K = 1 day or 1 month. At minimum, the paper should state this as an explicit assumption and identify what measurements would validate it; ideally it should include a small trace analysis or a concrete falsifiable prediction.
- [§1, numerical example] The headline example (transferring a 1 GB KV cache over a 10-Gbps link in ~0.8 s at ~$0.09 vs. recomputing in ~15 s at ~$0.47) assumes a fully utilized link and compares egress cost only. It does not account for storage cost over the reuse interval, lookup/discovery overhead in the proposed control plane, or the fact that the transferred cache must still be placed and managed. If the example is meant to motivate the vision, it should be hedged and broken down so the reader can see which costs are included and which are not.
minor comments (5)
- [§1, abstract] Typographical issues: "foran Internet" and "are evolve independently" should be fixed. The duplicate reference "[9, 17, 17, 43, 52, 56, 60]" should be cleaned.
- [§4.2, Figure 1] The phrase "Store Anywhere, Use EverywhereModel" is missing a space, and the figure should be referenced with a more explicit callout in the text explaining each control-plane component.
- [Eq. (1)] Symbol formatting is inconsistent: "KV MB" and "PrefillTFLOPs" contain spaces; consider using consistent math notation such as KV_MB and TFLOPs. Also, state explicitly that PeakFlopRate is in TFLOPs/hour so the denominator has units of dollars.
- [Table 1] Table 1 does not report the source values for DeviceCost_hr, TransferCost, GPUCost_hr, and PeakFlopRate. Since the table is central to the cost argument, these values should be listed or cited in a caption.
- [§5] Minor style issue: "new communication, policy and data-discovery abstractions ." has a stray space before the period.
Circularity Check
No significant circularity: the claimed 'new metrics' are algebraic restatements of a declared cost model, and the central vision is an external-data argument, not a self-referential derivation.
full rationale
The paper contains no fitted parameters, no trained model, and no empirical prediction that reduces to its own inputs. Equations (1) and (2) are algebra: K, KVMB, DeviceCost_hr, TransferCost, GPUCost_hr, PrefillTFLOPs, and PeakFlopRate are all declared inputs, and Eq. (2) is obtained by setting CacheCost/PrefillCost = 1 and cross-multiplying. The two 'new metrics' are exactly the two sides of that break-even condition, so labeling them metrics is a formal rearrangement rather than a circular derivation. The motivating comparison (1 GB over 10 Gbps in ~0.8 s at ~$0.09 vs ~15 s and ~$0.47 recompute) is arithmetic over external AWS prices and DeepSeek specifications and is independently checkable. Citations to the authors' own systems (LMCache [14], CacheBlend [66], METIS [54], DroidSpeak [41]) are used as examples of emerging KV-cache reuse mechanisms; the vision and cost model do not load-bearingly depend on the correctness of those systems. The main weakness is empirical: the paper assumes without measurements that cross-request, cross-boundary semantic reuse is frequent enough to make internet-scale KV-cache distribution cost-effective, and Section 3.3 itself concedes prefix caching is brittle. That is a missing-evidence/correctness concern, not circularity. The inconsistent use of K as a dimensionless frequency in Eq. (1) but as a time interval (1 day, 1 month) in Table 1 is a dimensional/correctness flaw, not a circular step.
Assumptions & free parameters
free parameters (3)
- K (KV Cache reuse frequency / retention time) =
1 day; 1 month (Table 1)
- Model-side ratio PrefillTFLOPs/KVMB =
11.410 (DeepSeek V4-Pro), 4.325 (DeepSeek V4-Flash)
- Reuse-horizon thresholds (hot/warm/cold) =
seconds / minutes / hours (Figure 1)
assumptions (4)
- domain assumption LLM workloads at global scale exhibit enough overlapping context reuse across requests, tenants, and regions to amortize cross-cloud KV Cache transfer.
- domain assumption Equation (1) is an adequate cost model: total cache cost equals storage cost plus transfer cost, with storage cost proportional to DeviceCosthr times K.
- domain assumption KV Cache state is portable across cloud providers, models, and engines via semantic or approximate reuse rather than exact prefix match.
- domain assumption Recompute cost is accurately represented by GPUCosthr times (PrefillTFLOPs / PeakFlopRate).
invented entities (2)
-
Global KV Cache control plane (KV Cache Directory, Lookup Service, Placement Planner, Prefetch Scheduler)
-
Internet for the KV Cache as an active content-distribution channel
Cite this review
Pith. "Pith review of An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age." pith.science (2026). https://pith.science/paper/EUTHUZPZ
@misc{pith2026260801526,
author = {Pith},
title = {Pith review of: An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age},
year = {2026},
howpublished = {\url{https://pith.science/paper/EUTHUZPZ}},
note = {Machine review of arXiv:2608.01526}
}
read the original abstract
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable context reuse from overlapping inputs, creating a major opportunity to store and reuse the contexts' KV Caches instead of recomputing them. However, model-side advances that shrink the KV Cache and system-side advances that reduce compute, storage, and transfer costs are evolve independently within legacy cloud boundaries. We argue that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters. The network becomes an active distribution channel; bandwidth, latency and pricing directly determines how the KV Cache should be managed. We propose a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system. In this view, KV Cache storage and recompute decisions are driven by model, infrastructure, and application metrics, to enable adaptive, content-driven decisions for minimizing latency and cost.
Figures
Reference graph
Works this paper leans on
-
[1]
Reyna Abhyankar, Qi Qi, and Yiying Zhang. 2026. OSWorld- Human: Benchmarking the Efficiency of Computer-Use Agents. arXiv:2506.16042 [cs.AI] https://arxiv.org/abs/2506.16042
arXiv 2026
-
[2]
Shubham Agarwal, Sai Sundaresan, Subrata Mitra, Debabrata Mahapa- tra, Archit Gupta, Rounak Sharma, Nirmal Joshua Kapu, Tong Yu, and Shiv Saini. 2025. Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation.Proceedings of the ACM on Manage- ment of Data3, 3 (2025), 136:1–136:28. https://doi.org/10.1145/3725273
doi:10.1145/3725273 2025
-
[3]
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ram- jee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, Santa Clara, CA, 117–134. https://www.useni...
work page 2024
-
[4]
Amazon Web Services. [n. d.]. Amazon EC2 Capacity Blocks for ML Pricing. https://aws.amazon.com/ec2/capacityblocks/pricing/. Accessed: 2026-07-16
work page 2026
-
[5]
Amazon Web Services. 2021. Overview of Data Transfer Costs for Common Architectures. https://aws.amazon.com/blogs/architect ure/overview-of-data-transfer-costs-for-common-architectures/. Accessed: 2026-07-09
work page 2021
-
[6]
Amazon Web Services. 2026. Accelerate LLM Model Loading and Increase Context Windows with GPUDirect on Amazon FSx for Lustre and TurboQuant. https://aws.amazon.com/blogs/machine-learning/a ccelerate-llm-model-loading-and-increase-context-windows-with- gpudirect-on-amazon-fsx-for-lustre-and-turboquant/. AWS Machine Learning Blog
work page 2026
-
[7]
Amazon Web Services. 2026. Amazon EC2 P6e UltraServers and P6 Instances. https://aws.amazon.com/ec2/instance-types/p6/. AWS product documentation
work page 2026
-
[8]
Amazon Web Services. 2026. Performance Guidelines for Amazon S3. https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimizi ng-performance-guidelines.html. Accessed: 2026-07-09
work page 2026
Show all 76 references
-
[9]
Anthropic. 2025. Prompt Caching with Claude. https://claude.com/b log/prompt-caching. Claude Blog
2025
-
[10]
Kyle Aubrey and Nick Stam. 2025. Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era. https://developer.nvidia.c om/blog/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai- factory-era/. NVIDIA Technical Blog
2025
-
[11]
Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai, Zhiding Yu, and Alexey Tumanov. 2025. RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression. InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Lear...
2025
- [13]
-
[14]
Yihua Cheng, Yuhan Liu, Jiayi Yao, Yuwei An, Xiaokun Chen, Shaoting Feng, Yuyang Huang, Samuel Shen, Kuntai Du, and Junchen Jiang
-
[15]
DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million- Token Context Intelligence. arXiv:2606.19348 [cs.CL] https://arxiv.or g/abs/2606.19348
2026
-
[16]
Florin Dobrian, Vyas Sekar, Asad Awan, Ion Stoica, Dilip Joseph, Aditya Ganjam, Jibin Zhan, and Hui Zhang. 2011. Understanding the Impact of Video Quality on User Engagement. InProceedings of the ACM SIGCOMM 2011 Conference (SIGCOMM ’11). Association for Com- puting Machinery,...
2011
-
[17]
Amr Elmeleegy and Akshatha Kamath. 2025. How to Reduce KV Cache Bottlenecks with NVIDIA Dynamo. https://developer.nvidia.com/b log/how-to-reduce-kv-cache-bottlenecks-with-nvidia-dynamo/. NVIDIA Technical Blog
2025
-
[18]
Zhiyuan Fang, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu, Fan Yu, and Zibin Zheng. 2025. Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline. InProceedings of the 30th ACM International Conference on Archi- tectural Support for P...
2025
-
[19]
Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, and Jie Wu. 2025. WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-Based Dynamic Scheduling. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association ...
2025
-
[20]
Aditya Ganjam, Faisal Siddiqui, Jibin Zhan, Xi Liu, Ion Stoica, Junchen Jiang, Vyas Sekar, and Hui Zhang. 2015. C3: Internet-Scale Control Plane for Video Quality Optimization. In12th USENIX Symposium on Networked Systems Design and Implementation (NSDI ’15). USENIX Associatio...
2015
-
[21]
Google Cloud. 2026. Bucket Locations. https://cloud.google.com/sto rage/docs/locations. Accessed: 2026-07-09
2026
-
[22]
Google Cloud. 2026. Network Pricing. https://cloud.google.com/vpc /network-pricing. Accessed: 2026-07-09
2026
-
[23]
Ke Hong, Xiuhong Li, Lufang Chen, Qiuli Mao, Guohao Dai, Xuefei Ning, Shengen Yan, Yun Liang, and Yu Wang. 2025. SOLA: Optimizing SLO Attainment for Large Language Model Serving with State-Aware Scheduling. InProceedings of Machine Learning and Systems, Vol. 7. https://proceed...
2025
-
[24]
Junhao Hu, Wenrui Huang, Weidong Wang, Haoyi Wang, Tiancheng Hu, Zhang Qin, Hao Feng, Xusheng Chen, Yizhou Shan, and Tao Xie
-
[25]
Mengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu, Wenqi Shao, and Ping Luo. 2025. HiAgent: Hierarchical Working Memory Man- agement for Solving Long-Horizon Agent Tasks with Large Language Model. InProceedings of the 63rd Annual Meeting of the Association for Computational Lin...
2025 doi
-
[26]
Language Models
EPIC: Efficient Position-Independent Caching for Serving Large 7 Ray et al. Language Models. InProceedings of the 42nd International Confer- ence on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Juli...
-
[27]
Te-Yuan Huang, Ramesh Johari, Nick McKeown, Matthew Trunnell, and Mark Watson. 2014. A Buffer-Based Approach to Rate Adaptation: Evidence from a Large Video Streaming Service. InProceedings of the 2014 ACM Conference on SIGCOMM (SIGCOMM ’14). Association for Computing Machiner...
2014 doi
-
[29]
Patil, Joseph E
Paras Jain, Sam Kumar, Sarah Wooders, Shishir G. Patil, Joseph E. Gonzalez, and Ion Stoica. 2022. Skyplane: Optimizing Transfer Cost and Throughput Using Cloud-Aware Overlays.arXiv preprint arXiv:2210.07259(2022). arXiv:2210.07259 [cs.DC] https://arxiv.org/ abs/2210.07259
2022 arXiv
-
[30]
Internet Engineering Task Force. 2026. Introduction to the IETF. https://www.ietf.org/about/introduction/. Accessed July 2026
2026
-
[31]
Youhe Jiang, Fangcheng Fu, Xiaozhe Yao, Taiyi Wang, Bin Cui, Ana Klimovic, and Eiko Yoneki. 2025. ThunderServe: High-Performance and Cost-Efficient LLM Serving in Cloud Environments. InProceedings of Machine Learning and Systems, Vol. 7. https://proceedings.mlsys.or g/paper_fi...
2025
-
[32]
Junchen Jiang, Vyas Sekar, and Hui Zhang. 2012. Improving Fairness, Efficiency, and Stability in HTTP-Based Adaptive Video Streaming with FESTIVE. InProceedings of the 8th International Conference on Emerging Networking Experiments and Technologies (CoNEXT ’12). As- sociation ...
2012
-
[33]
Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Shufan Liu, Xuanzhe Liu, and Xin Jin. 2026. RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.ACM Transactions on Computer Systems44, 1 (2026), 2:1–2:27. https://doi.org/10.1145/3768628
2026 doi
-
[34]
Shibo Jie, Yehui Tang, Kai Han, Yitong Li, Duyu Tang, Zhi-Hong Deng, and Yunhe Wang. 2025. Mixture of Lookup Experts. InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Daniel ...
2025
-
[35]
Lee, Sangdoo Yun, and Hyun Oh Song
Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, and Hyun Oh Song. 2025. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction. InAdvances in Neural Information Processing Systems, Vol. 38. https://proceedings.neurips.cc /paper_files/paper/2025...
2025
-
[36]
Keisuke Kamahori, Tian Tang, Yile Gu, Kan Zhu, and Baris Kasikci
-
[37]
InThe Thirteenth International Conference on Learn- ing Representations
Fiddler: CPU–GPU Orchestration for Fast Inference of Mixture- of-Experts Models. InThe Thirteenth International Conference on Learn- ing Representations. https://proceedings.iclr.cc/paper_files/p aper/2025/hash/8cd1ce03ea58b3d7dfd809e4d42f08ea-Abstract- Conference.html
2025
-
[38]
Dong Liu and Yanxuan Yu. 2026. CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving. InProceedings of the 2026 ACM/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA ’26). Association for Computing Machinery, 56–66. https://doi.or...
2026
-
[39]
Xunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma, and Xun Zhou. 2025. FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=OfjIlb elrT
2025
- [40]
-
[41]
Yuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng, Zhuohan Gu, Kuntai Du, Hanchen Li, Yihua Cheng, Junchen Jiang, Shan Lu, Madan Musuvathi, and Esha Choukse. 2026. DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants. InProceedings of the 23rd USENIX Symposium on Netw...
2026
-
[42]
Gonzalez, Danny Harnik, and Ion Stoica
Shu Liu, Xiangxi Mo, Moshik Hershcovitch, Henric Zhang, Audrey Cheng, Guy Girmonsky, Gil Vernik, Michael Factor, Tiemo Bang, Soujanya Ponnapalli, Natacha Crooks, Joseph E. Gonzalez, Danny Harnik, and Ion Stoica. 2025. SkyStore: Cost-Optimized Object Stor- age Across Regions an...
2025 arXiv
-
[43]
Xiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li, Liuyue, Bo Li, Xum- ing Hu, and Xiaowen Chu. 2025. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference. In Advances in Neural Information Processing Systems, Vol. 38. https: //proceedings.ne...
2025
-
[44]
Ziming Mao, Tian Xia, Zhanghao Wu, Wei-Lin Chiang, Tyler Griggs, Romil Bhardwaj, Zongheng Yang, Scott Shenker, and Ion Stoica. 2025. SkyServe: Serving AI Models across Regions and Clouds with Spot Instances. InProceedings of the Twentieth European Conference on Com- puter Syst...
2025
-
[45]
llm-d Contributors. 2026. llm-d-kv-cache: Distributed KV Cache Sched- uling and Offloading Libraries. https://github.com/llm-d/llm-d-kv- cache. Accessed: 2026-07-08
2026
-
[46]
LMCache. 2026. Stop Calling It KV Cache: It’s Something Much Bigger. https://blog.lmcache.ai/en/2026/04/28/stop-calling-it-kv-cache-its- something-much-bigger/. LMCache Blog
2026
-
[47]
Microsoft Azure. 2026. Architecture Best Practices for Azure Files. https://learn.microsoft.com/en-us/azure/well-architected/service- 8 An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age guides/azure-files. Accessed: 2026-07-09
2026
-
[48]
Micron Technology. 2025. Micron Ships HBM4 to Key Customers to Power Next-Gen AI Platforms. https://investors.micron.com/news- releases/news-release-details/micron-ships-hbm4-key-customers- power-next-gen-ai-platforms. Micron press release
2025
-
[49]
Micron Technology. 2025. Micron Unveils Portfolio of Industry-First SSDs to Power the AI Revolution. https://investors.micron.com/news- releases/news-release-details/micron-unveils-portfolio-industry- first-ssds-power-ai-revolution. Micron press release
2025
-
[50]
NVIDIA. 2026. TensorRT-LLM: KV Cache System. https://github.com /NVIDIA/TensorRT-LLM/blob/main/docs/source/features/kvcache .md. Accessed: 2026-07-08
2026
-
[51]
NVIDIA. 2026. NVIDIA Dynamo: A Datacenter-Scale Distributed Inference Framework. https://docs.nvidia.com/dynamo/getting- started/introduction. Accessed: 2026-07-08
2026
-
[52]
NVIDIA. 2026. NVIDIA Kicks Off the Next Generation of AI With Rubin. https://nvidianews.nvidia.com/news/rubin-platform-ai- supercomputer. NVIDIA Newsroom
2026
-
[53]
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading More Storage for Less Computation — A KVCache-Centric Architecture for Serving LLM Chatbot. In23rd USENIX Conference on File and Storage Tec...
2025
-
[54]
OpenAI. 2026. ChatGPT Agent. https://help.openai.com/en/articles/ 11752874-chatgpt-agent. OpenAI Help Center
2026
-
[55]
OpenAI. 2026. Prompt Caching 201. https://developers.openai.co m/cookbook/examples/prompt_caching_201. OpenAI Developers Cookbook
2026
-
[56]
Kai Shen and Barak Epstein. 2025. Reducing TCO for AI Inferencing with External KV Cache on Managed Lustre. https://cloud.google.c om/blog/products/storage-data-transfer/choosing-google-cloud- managed-lustre-for-your-external-kv-cache. Google Cloud Blog
2025
-
[57]
Siddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du, Shaoting Feng, Ganesh Ananthanarayanan, Ravi Netravali, and Junchen Jiang. 2025. METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. ...
2025
-
[58]
Chaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi, Yongji Wu, Jialin Li, and Cheng Li. 2026. Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM Workloads. In23rd USENIX Symposium on Networked Systems Design and Imple- mentation (NSDI 26)...
2026
-
[59]
The AIBrix Team, Jiaxin Shan, Varun Gupta, Le Xu, Haiyang Shi, Jingyuan Zhang, Ning Wang, Linhui Xu, Rong Kang, Tongping Liu, Yifei Zhang, Yiqing Zhu, Shuowei Jin, Gangmuk Lim, Bin- bin Chen, Zuzhi Chen, Xiao Liu, Xin Chen, Kante Yin, Chak-Pong Chung, Chenyu Jiang, Yicheng Lu,...
2025 arXiv
-
[60]
Qidong Su, Wei Zhao, Xin Li, Muralidhar Andoorveedu, Chenhao Jiang, Zhanda Zhu, Kevin Song, Christina Giannoula, and Gennady Pekhimenko. 2025. Seesaw: High-Throughput LLM Inference via Model Re-Sharding. InProceedings of Machine Learning and Systems, Vol. 7. https://proceeding...
2025
-
[61]
Weiwei Sun, Miao Lu, Zhan Ling, Kang Liu, Xuesong Yao, Yiming Yang, and Jiecao Chen. 2026. Scaling Long-Horizon Agent via Con- text Folding. InProceedings of the 43rd International Conference on Machine Learning. https://openreview.net/forum?id=JaLXQnA2wi Forthcoming
2026
-
[62]
Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. 2025. DuoAttention: Ef- ficient Long-Context LLM Inference with Retrieval and Streaming Heads. InThe Thirteenth International Conference on Learning Repre- sentations. https...
2025
-
[63]
vLLM Project. 2026. Automatic Prefix Caching. https://docs.vllm.ai/e n/latest/features/automatic_prefix_caching/. vLLM Documentation
2026
-
[64]
vLLM Project. 2026. vLLM: Automatic Prefix Caching. https://docs.vll m.ai/en/stable/design/prefix_caching/. Accessed: 2026-07-08
2026
-
[65]
Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Frank Sifei Luan, Gautam Mittal, Scott Shenker, and Ion Stoica. 2023. SkyPilot: An Intercloud Broker for Sky Computing. InProceedings of the 20th USENIX Sym- posium on Networke...
2023
-
[66]
Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. 2025. A-Mem: Agentic Memory for LLM Agents. InAdvances in Neural Information Processing Systems, Vol. 38. https://proceedings. neurips.cc/paper_files/paper/2025/hash/19909c36f51abc4856b4560 aff3d36d6-A...
2025
-
[67]
Lijie Yang, Zhihao Zhang, Zhuofu Chen, Zikun Li, and Zhihao Jia
-
[68]
InThe Thirteenth International Conference on Learning Representations
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention. InThe Thirteenth International Conference on Learning Representations. https://proceedings.iclr.cc/paper_fil es/paper/2025/hash/11440c427f0f76f191ac06b50d7a2517-Abstract- Conference.html
2025
-
[69]
Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and Hao Zhou. 2026. MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-Based Memory Agent. InThe Fourteenth International Conference on L...
2026
-
[70]
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. 2025. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowl- edge Fusion. InProceedings of the Twentieth European Conference on Computer Systems...
2025
-
[71]
Xiaoqi Yin, Abhishek Jindal, Vyas Sekar, and Bruno Sinopoli. 2015. A Control-Theoretic Approach for Dynamic Adaptive Video Streaming over HTTP. InProceedings of the 2015 ACM Conference on Special Interest Group on Data Communication (SIGCOMM ’15). Association for Computing Mac...
2015 doi
-
[72]
Noh, and Jongry- ool Kim
Dongha Yoon, Younghoon Min, Hoshik Kim, Sam H. Noh, and Jongry- ool Kim. 2025. TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale.arXiv preprint arXiv:2512.18194 (Dec. 2025). https://doi.org/10.48550/arXiv.2512.18194 arXiv:2512.18194 [cs.DC]
2025 doi
-
[73]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. In18th USENIX Symposium on Operating Systems Design and Implementation (O...
2024
-
[74]
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. 2025. Native Sparse Attention: Hardware- Aligned and Natively Trainable Sparse A...
2025
-
[75]
Amir Zandieh, Majid Daliri, Majid Hadian, and Vahab Mirrokni. 2026. TurboQuant: Online Vector Quantization with Near-Optimal Distor- tion Rate. InThe Fourteenth International Conference on Learning Rep- resentations. https://openreview.net/forum?id=tO3ASKZlok 9 Ray et al
2026
-
[76]
Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Meng- meng Ji, Hanchen Li, Urmish Thakker, James Zou, and Kunle Oluko- tun. 2026. Agentic Context Engineering: Evolving Contexts for Self- Improving Language Models...
2026
-
[210]
https://www.usenix.org/conference/osdi24/presentation/zhong- yinmin 10
-
[2025]
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference.arXiv preprint arXiv:2510.09665(Oct. 2025). https: //doi.org/10.48550/arXiv.2510.09665 arXiv:2510.09665 [cs.DC]
2025 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.