REVIEW 3 major objections 6 minor 87 references
Selective transfer of only the important KV cache blocks can cut time-to-second-token by up to 4.3× on bandwidth-limited cloud GPUs without hurting later decoding or accuracy.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 16:18 UTC pith:AMO72SYU
load-bearing objection Solid systems fix for a real cloud P/D bottleneck: selective KV transfer with three paths delivers the claimed TTST cut; TBT parity rests more on the parallel/speculative lanes than on the profiling bet. the 3 major comments →
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By ranking KV blocks according to how often they are selected across a calibration set, a system can transfer only a small, positionally stable subset during prefill and still keep subsequent decoding both fast and accurate; the missing pieces are obtained by overlapping remote fetches with local memory loads and by background speculative transfer.
What carries the argument
Three coordinated transfer paths—profile-based proactive push of the top-Kr universally important blocks, parallel on-demand fetch that splits the KV index into local and remote lanes, and speculative background delivery of the remainder during attention idle time.
Load-bearing premise
Important cache positions stay roughly the same across prompts of similar length, so an offline frequency matrix measured on one calibration set remains a good ranking for unseen online requests.
What would settle it
Re-run the MultiFieldQA / GovReport experiments after replacing the profile-derived importance matrix with a random or purely sequential ranking of the same number of blocks; if the on-demand ratio and time-to-second-token no longer improve, positional similarity does not hold.
If this is right
- Self-hosted disaggregated serving becomes practical on 15–32 Gbps cloud instances without needing 100+ Gbps fabrics.
- Any existing dynamic sparse-attention algorithm that emits a KV-index matrix can be plugged in without changing its selection logic.
- Under light network load the same analytical model automatically falls back to full transfer, preserving bandwidth for other traffic.
- Quantization schemes remain complementary; the two techniques can be stacked.
Where Pith is reading between the lines
- If positional similarity weakens on highly domain-shifted traffic, a cheap online re-profiling pass every few thousand requests would restore the gains without redesigning the three paths.
- The same importance matrix could also guide which blocks to keep in limited GPU HBM versus host memory, turning the transfer engine into a unified cache-placement policy.
- Because the speculative path already uses idle NIC cycles, the same machinery could ship prefix-cache hits or multi-turn conversation state with almost no extra latency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SmartGen addresses stage-transition stall in P/D-disaggregated LLM inference on bandwidth-limited cloud GPUs by selectively transferring KV cache rather than shipping it in full. It partitions KV entries into universally important, context-dependent, and less-important classes and routes them over three paths: (1) profile-based proactive RDMA WRITE of top-Kr blocks ranked by an offline importance matrix I (Eq. 1–2, §4.1); (2) parallel on-demand fetch that splits the decode-side KV index via a GPU-resident 0/1/2 mask and overlaps remote GDR with local H2D (§4.2); (3) speculative background delivery of the remainder during attention-idle NIC/CPU windows at a fixed 10% ratio (§4.3). Evaluated on Llama-3.1-8B, Qwen3-8B/14B, Gemma-3-12B, and Phi-4-14B against full transfer, sequential partial transfer, and HACK (int2), the system reports up to 4.3× lower TTST with near-parity TBT after a few tokens and LongBench accuracy comparable to full-cache InfiniGen/HATA baselines.
Significance. The work targets a practical deployment gap: self-hosted disaggregated serving on rented instances whose 10–35 Gbps NICs cannot hide full KV transfer behind prefill. The three-path design is concrete and systems-complete (mask protocol, doorbell-batched WRITE, scatter RECV, analytical Kr). Strengths include multi-model/multi-workload wall-clock results, factor ablation (Fig. 13), batch/bandwidth/sequence/M/selection-ratio sweeps, and explicit accuracy curves versus full cache and HACK (Figs. 9, 17). Orthogonality to quantization (HACK) is correctly claimed. If the positional-similarity assumption and TBT-parity results hold under broader workload shift, the paper offers a usable blueprint for bandwidth-constrained P/D stacks.
major comments (3)
- [§4.1, Fig. 5, Eq. 1, Fig. 15] §4.1, Fig. 5, Eq. 1: The comparable-TBT half of the central claim rests on positional similarity—that an importance matrix I profiled on held-out 2WikiMultihopQA ranks hot blocks for the four evaluation LongBench tasks. Fig. 15 shows profile beating sequential/random on MultiFieldQA, and Fig. 13 attributes only 1.03–1.1× of the TBT recovery to profiling (parallel + speculative do most of the work). The manuscript does not stress-test domain or template shift outside LongBench (e.g., pure code, multi-turn chat, or RAG with retrieved passages whose hot positions differ). If online top-Kr is cold, on-demand ratio rises and TBT parity with full transfer can fail for the first tens of tokens even while TTST stays strong. A cross-domain or cross-template ablation, or an online I refresh experiment, is needed to bound this risk.
- [§3, §5.1, Eq. 2] §3 and §5.1: All main results assume a 75% host-memory prefix-cache hit ratio, which shortens effective prefill work and therefore sets Kr via Eq. 2 (Tp/Tt). No sensitivity sweep over hit ratio appears in §5.4. Because Kr directly controls how much volume is removed from the stage-transition path, the headline TTST gains (and the claim that selective transfer fully overlaps prefill) are conditioned on this free parameter. A hit-ratio sweep (e.g., 0%, 50%, 75%, 90%) on at least one model/workload would show whether the 4.3× result is robust when prefix reuse is weaker.
- [§4.3, Fig. 14] §4.3 and Fig. 14: Speculative transfer is capped at a fixed 10% of remaining blocks per idle window and is justified by a single-model MultifieldQA trace. Under higher decode concurrency or when multiple prefill nodes contend for the same decode NIC, the assumed attention-idle NIC/CPU slack shrinks; the paper does not measure interference or TTST/TBT under multi-request decode batches beyond the reported batch-size sweep. Clarifying the concurrency regime in which the 10% ratio remains non-intrusive is load-bearing for the “comparable subsequent decoding” claim once speculative fill is the mechanism that drives TBT to the full-transfer floor.
minor comments (6)
- [Fig. 9] Fig. 9 y-axis is average CTL (s) but the narrative mixes CTL, TTST, and TBT; a short caption note defining CTL = cumulative latency up to token t would help readers parse the early spike vs later slope.
- [§4.1, Eq. 2] Eq. 2 writes Kr = M·(L−1)·Tp/Tt and then clips to [2M, LM]. The text should state explicitly whether Tp/Tt are measured per-layer averages or end-to-end wall-clock, and whether metadata tensors (InfiniGen partial keys, etc.) are folded into Tt.
- [§4.4] §4.4 mentions fallback to full transfer under low load via Eq. 2, but no experiment shows the switch point. A one-sentence result or appendix plot would make the adaptivity claim concrete.
- [Table 1, Fig. 12] Table 1 lists 32/25/15 Gbps; Fig. 12 uses 2xlarge/4xlarge/2x.8xlarge. Align instance names and bandwidth labels across table and figure to avoid cross-referencing friction.
- [§2.1, Fig. 11, §5.1] Typos/notation: “so f tmax” spacing in §2.1; “infinite.” vs “Infini.” in Fig. 11 legend; “SmartGen(Infini./HATA)” formatting is inconsistent in §5.1 vs figure legends.
- [§6.3] Related work §6.3 correctly notes orthogonality to HACK; a brief quantitative note on whether SmartGen’s selective paths compose with 2-bit (or 4-bit) KV would strengthen the “complementary” claim without requiring a full joint implementation.
Circularity Check
No circularity: empirical systems paper; TTST/TBT/accuracy are measured wall-clock outcomes, not identities forced by fitted inputs or self-citation.
full rationale
SmartGen’s central claims (up to 4.3× TTST vs full transfer; comparable TBT and LongBench accuracy) are established by end-to-end runs against Full/Partial/HACK baselines on cloud L20/V100S instances, not by algebraic identities. Eq. 1 defines block importance I_{l,m} as offline selection frequency on a held-out calibration split (LongBench 2WikiMultihopQA); Eq. 2 sets the proactive budget K_r = M·(L−1)·T_p/T_t from profiled prefill vs transfer times—an engineering overlap bound, not a fitted prediction of the headline speedup. Decode-side InfiniGen/HATA still select dynamically and parallel on-demand fetches misses, so accuracy is checked against full-cache ground truth rather than assumed from the profile. Factor analysis (Fig. 13) and strategy ablations (Fig. 15) treat profile quality as an empirical variable that can fail. Citations (InfiniGen, HATA, HACK, Mooncake, etc.) are external prior art used as components or baselines, not load-bearing uniqueness theorems by the same authors. No step reduces the claimed result to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (5)
- M (KV blocks per layer) =
1000
- speculative_ratio =
10%
- prefix_cache_hit_ratio =
75%
- InfiniGen/HATA selection hyperparameters =
alpha=5, partial=0.3, max_sel=20%
- Kr overlap budget =
M*(L-1)*Tp/Tt (clipped)
axioms (5)
- domain assumption Dynamic KV sparsity: attending to a small top-k subset of KV entries preserves task accuracy for the evaluated long-context benchmarks.
- domain assumption Positional similarity: important KV entries concentrate in stable sequence regions across prompts of similar length, so a calibration dataset predicts online block importance.
- domain assumption Permutation of selected KV entries along the sequence axis does not change attention output because positional information is already baked into K/V.
- domain assumption RDMA NICs provide in-order delivery so doorbell-batched WRITE of KV block then mask is safe without extra fencing.
- standard math Standard transformer attention and P/D disaggregation cost model (prefill compute-bound, decode memory-bound, KV size ∝ layers×batch×seq×dim).
invented entities (3)
-
Three-class KV taxonomy (universally important / context-dependent / less important) with matching transfer paths
independent evidence
-
GPU-resident KV mask matrix (values 0/1/2) updated by remote RDMA WRITE
independent evidence
-
Importance matrix I and block score Il,m (Eq. 1)
independent evidence
read the original abstract
Disaggregating the prefill and decoding stages of large language model (LLM) inference into two separate sets of nodes is widely adopted in today's LLM serving systems. However, such an architecture poses significant challenges for self-hosted LLM deployments on rented cloud instances, since transferring enormous key-value (KV) caches between disaggregated nodes can easily saturate the limited inter-node network bandwidth. In this paper, we propose to mitigate the network bottleneck by selectively transferring essential KV cache entries across the two stages. There are two challenges to achieve selective KV cache transfer, i.e., accurate KV selection during the prefill stage, and efficient KV fetching during the decoding stage. To address these challenges, we design SmartGen, a KV cache transfer engine that allows seamless disaggregated LLM inference with three data transfer paths. Specifically, we leverage 1) a profile-based proactive transfer path to identify and push essential KV cache entries to the decoding node during the prefill stage, 2) a parallel on-demand transfer path to simultaneously fetch remote and local KV cache entries during the decoding stage, and 3) a speculative transfer path to finally deliver all KV caches to the decoding node. Experimental results show that SmartGen reduces time-to-second-token by up to 4.3x compared with the typical full KV cache transfer approach while offering comparable subsequent decoding performance and accuracy.
Figures
Reference graph
Works this paper leans on
-
[1]
Marah I Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat S. Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Java- heripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, Vibhav Vineet, Yue Wu, Sa...
Pith/arXiv arXiv 2025
-
[2]
Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In18th USENIX Symposium on Oper- ating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 117–134. USENIX Ass...
2024
-
[3]
Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025
InfiniBand Trade Association. Infiniband architecture specification volume 1 release 1.8.https://www.infini bandta.org/ibta-specification, Accessed: 2025
2025
-
[4]
LongBench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. LongBench: A bilingual, multitask benchmark for long context understanding. InProceedings of the 62nd An- nual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), AC...
2024
-
[5]
TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing
Junyi Chen, Chuheng Du, Renyuan Liu, Shuochao Yao, Dingtian Yan, Jiang Liao, Shengzhong Liu, Fan Wu, and Guihai Chen. TokenFlow: Responsive LLM text stream- ing serving under request burst via preemptive schedul- ing. InProceedings of the 21st European Conference on Computer Systems, EuroSys 2026, 2026. ACM, 2025
2026
-
[6]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brock- man, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavari...
Pith/arXiv arXiv 2021
-
[7]
Retroinfer: A vector storage engine for scalable long-context LLM inference
Yaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang, Chengruidong Zhang, Jing Liu, Jingjia Luo, Di Liu, Huiqiang Jiang, Qi Chen, Bailu Ding, Xiao Yan, Jiawei Jiang, Chen Chen, Mingxing Zhang, Cheng Li, Yuqing Yang, Fan Yang, and Mao Yang. Retroinfer: A vector storage engine for scalable long-context LLM inference. Proc. VLDB Endow., 19(5):1016–1031, 2026
2026
-
[8]
Elastic GPU service instance families
Alibaba Cloud. Elastic GPU service instance families. https://www.alibabacloud.com/help/en/ecs/user-gui de/gpu-accelerated-compute-optimized-and-vgpu-a ccelerated-instance-families-1, Accessed: 2025
2025
-
[9]
eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025
Alibaba Cloud. eRDMA.https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025
2025
-
[10]
A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025
Google Cloud. A2 ultra machine types.https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025
2025
-
[11]
Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025
Tencent Cloud. Computing instance.https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025
2025
-
[12]
NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025
NVIDIA Corporation. NVIDIA dynamo platform.https: //developer.nvidia.com/dynamo, Accessed: 2025
2025
-
[13]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memory- efficient exact attention with io-awareness. InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022
2022
-
[14]
Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025
DeepSeek-AI. Deepseek-v3.2-exp: Boosting long- context efficiency with deepseek sparse attention.https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025. 13
2025
-
[15]
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingx- uan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Er- hang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, ...
Pith/arXiv arXiv 2024
-
[16]
Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications
Kuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng, Qing Lan, Hejian Sang, Yihua Cheng, Jiayi Yao, Xiaox- uan Liu, Yifan Qiao, Ion Stoica, and Junchen Jiang. Pre- fillOnly: An inference engine for prefill-only workloads in large language model applications. InProceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel...
2025
-
[17]
The design and operation of CloudLab
Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuang-Ching Wang, Glenn Ricart, Larry Landwe- ber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. The design and operation of CloudLab. InProceedings of the 2019 ...
2019
-
[18]
Graham, Artem Y
Ana Gainaru, Richard L. Graham, Artem Y. Polyakov, and Gilad Shainer. Using infiniband hardware gather- scatter capabilities to optimize MPI all-to-all. InPro- ceedings of the 23rd European MPI Users’ Group Meeting, EuroMPI 2016, Edinburgh, United Kingdom, September 25-28, 2016, pages 167–179. ACM, 2016
2016
-
[19]
Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. Cost-efficient large language model serving for multi-turn conversations with CachedAtten- tion. InProceedings of the 2024 USENIX Annual Technical Conference, USENIX ATC 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 111–126. USENIX A...
2024
-
[20]
Fast state restoration in LLM serving with HCache
Shiwei Gao, Youmin Chen, and Jiwu Shu. Fast state restoration in LLM serving with HCache. InProceed- ings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 128–143. ACM, 2025
2025
-
[21]
Weaver: Efficient multi-llm serving with attention offloading
Shiwei Gao, Qing Wang, Shaoxun Zeng, Youyou Lu, and Jiwu Shu. Weaver: Efficient multi-llm serving with attention offloading. InProceedings of the 2025 USENIX Annual Technical Conference, USENIX ATC 2025, Boston, MA, USA, July 7-9, 2025, pages 587–595. USENIX Asso- ciation, 2025
2025
-
[22]
Hybrid multi- document summarization using pre-trained language models.Expert Syst
Alireza Ghadimi and Hamid Beigy. Hybrid multi- document summarization using pre-trained language models.Expert Syst. Appl., 192:116292, 2022
2022
-
[23]
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Alek- sander Wawer. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization.CoRR, abs/1911.12237, 2019
Pith/arXiv arXiv 1911
-
[24]
HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference
Ping Gong, Jiawei Yi, Shengnan Wang, Juncheng Zhang, Zewen Jin, Ouxiang Zhou, Ruibo Liu, Guanbin Xu, Youhui Bai, Bowen Ye, Kun Yuan, Tong Yang, Gong Zhang, Renhai Chen, Feng Wu, and Cheng Li. HATA: trainable and hardware-efficient hash-aware top-k at- tention for scalable large model inference. InFindings of the Association for Computational Linguistics, ...
2025
-
[25]
Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization. InProceedings of the 50th Annual International Sympo- sium on Computer Architecture, ISCA 2023, Orlando, FL, USA, June 17-21, 2023, pages 3:1–3:15. ACM, 2023
2023
-
[26]
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Ju- lian J. McAuley. LongCoder: A long-range pre-trained language model for code completion. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 ofProceed- ings of Machine Learning Research, pages 12098–12107. PMLR, 2023. 14
2023
-
[27]
OmniKV: Dynamic context selection for efficient long-context LLMs
Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, and Sheng Guo. OmniKV: Dynamic context selection for efficient long-context LLMs. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025
2025
-
[28]
WaferLLM: Large language model inference at wafer scale
Congjie He, Yeqi Huang, Pei Mu, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang, and Luo Mai. WaferLLM: Large language model inference at wafer scale. In19th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2025, Boston, MA, USA, July 7-9, 2025, pages 257–273. USENIX Association, 2025
2025
-
[29]
Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps. InProceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 6609–6625. International Committee on Computational Linguis...
2020
-
[30]
Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. KVQuant: Towards 10 million con- text length LLM inference with KV cache quantization. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Process- ing Systems 2024, NeurIPS 2024, Vancouver, BC,...
2024
-
[31]
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan. Inference without interference: Disaggregate LLM in- ference for mixed downstream workloads.CoRR, abs/2401.11181, 2024
Pith/arXiv arXiv 2024
-
[32]
Efficient attentions for long document summarization
Luyang Huang, Shuyang Cao, Nikolaus Nova Parulian, Heng Ji, and Lu Wang. Efficient attentions for long document summarization. InProceedings of the 2021 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 1419–1436. Association for Computat...
2021
-
[33]
Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar
Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Si- mon Peter, Ramachandran Ramjee, and Ashish Panwar. POD-Attention: Unlocking full prefill-decode overlap for faster LLM inference. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Vol- ume 2, ASPLOS 2025, Rotterdam, Netherlan...
2025
-
[34]
Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization
Minsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi, Hunjong Lee, Junsoo Kim, Joo-Young Kim, and Jongse Park. Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization. In Proceedings of the 52nd Annual International Symposium on Computer Architecture, ISCA 2025, Tokyo, Japan, June 21-25, 2025, pages 482–497. ACM, 2025
2025
-
[35]
Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains
Abhishek Vijaya Kumar, Gianni Antichi, and Rachee Singh. Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Sys- tems, Volume 2, ASPLOS 2025, Rotterdam, Netherlands, 30 March 2025 - 3 April 2025, pages 48–62. ACM, 2025
2025
-
[36]
Efficient memory management for large language model serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. InProceedings of the 29th Symposium on Operating Sys- tems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, pages 611–626. ACM, 2023
2023
-
[37]
InfiniGen: Efficient generative inference of large language models with dynamic KV cache management
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. InfiniGen: Efficient generative inference of large language models with dynamic KV cache management. In18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 155–172. USENIX Association, 2024
2024
-
[38]
ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression
Guangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang, and Minyi Guo. ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression. pages 1–7, 2025
2025
-
[39]
Cachegen: KV cache compression and streaming for fast large language model serving
Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, and Junchen Jiang. Cachegen: KV cache compression and streaming for fast large language model serving. InProceedings of the ACM SIGCOMM 2024 Conference, ACM SIGCOMM 2024, Sydney...
2024
-
[40]
Helix: Serving large language models over heterogeneous GPUs and network via max-flow
Yixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang, Zhihao Jia, and Rashmi Vinayak. Helix: Serving large language models over heterogeneous GPUs and network via max-flow. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The Netherlands, 30...
2025
-
[41]
Microsoft. MInference: Million-tokens prompt infer- ence for long-context llms.https://www.microsoft.co m/en-us/research/project/minference-million-token s-prompt-inference-for-long-context-llms, Accessed: 2025
2025
-
[42]
Heterogeneity-aware cluster scheduling policies for deep learning workloads
Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia. Heterogeneity-aware cluster scheduling policies for deep learning workloads. In14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020, Virtual Event, November 4-6, 2020, pages 481–498. USENIX Association, 2020
2020
-
[43]
GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025
OpenAI. GPT-5 is here.https://openai.com/gpt-5, Accessed: 2025
2025
-
[44]
InstAttention: In-storage attention offloading for cost-effective long-context LLM inference
Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, and Jie Zhang. InstAttention: In-storage attention offloading for cost-effective long-context LLM inference. InIEEE International Symposium on High Performance Computer Architecture, HPCA 2025, Las Vegas, NV, USA, March 1-5, 2025, pages 1510–1525. IEEE, 2025
2025
-
[45]
Splitwise: Efficient generative LLM inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient generative LLM inference using phase splitting. In51st ACM/IEEE Annual International Symposium on Computer Architecture, ISCA 2024, Buenos Aires, Argentina, June 29 - July 3, 2024, pages 118–132. IEEE, 2024
2024
-
[46]
vAttention: Dynamic memory management for serving llms with- out PagedAttention
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ra- machandran Ramjee, and Ashish Panwar. vAttention: Dynamic memory management for serving llms with- out PagedAttention. InProceedings of the 30th ACM International Conference on Architectural Support for Pro- gramming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The Netherlands, 30 March ...
2025
-
[47]
Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot. In23rd USENIX Conference on File and Storage Technologies, FAST 2025, Santa Clara, CA, February 25-27, 2025, pages 155–170. USENIX As-...
2025
-
[48]
Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather
Deepti Raghavan, Philip Alexander Levis, Matei Za- haria, and Irene Zhang. Breakfast of champions: to- wards zero-copy serialization with NIC scatter-gather. InHotOS ’21: Workshop on Hot Topics in Operating Sys- tems, Ann Arbor, Michigan, USA, June, 1-3, 2021, pages 199–205. ACM, 2021
2021
-
[49]
Code Llama: Open foundation models for code.CoRR, abs/2308.12950, 2023
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wen- han Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Tho...
Pith/arXiv arXiv 2023
-
[50]
Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025
Amazon Web Services. Partner success with AWS.http s://aws.amazon.com/partners/success, Accessed: 2025
2025
-
[51]
Recommended GPU instances
Amazon Web Services. Recommended GPU instances. https://docs.aws.amazon.com/dlami/latest/devguide/ gpu.html, Accessed: 2025
2025
-
[52]
Gonzalez, and Ion Stoica
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuo- han Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. Fairness in serving large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 965–988. USENIX Association, 2024
2024
-
[53]
FlexGen: High-throughput generative inference of large language models with a single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. FlexGen: High-throughput generative inference of large language models with a single GPU. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 ofProceedings of Machin...
2023
-
[54]
Maguire Jr., and Dejan Kostic
Xiangyu Shi, Marco Chiesa, Gerald Q. Maguire Jr., and Dejan Kostic. KVComm: Enabling efficient LLM com- munication through selective KV sharing. 2026
2026
-
[55]
Sudipta Saha Shubha, Haiying Shen, and Anand P. Iyer. USHER: holistic interference avoidance for resource optimized ML inference. In18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 947–
2024
-
[56]
PowerInfer: Fast large language model serving with a consumer-grade GPU
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen. PowerInfer: Fast large language model serving with a consumer-grade GPU. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP 2024, Austin, TX, USA, November 4-6, 2024, pages 590–606. ACM, 2024
2024
-
[57]
Preble: Efficient distributed prompt scheduling for LLM serving
Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dong- ming Li, and Yiying Zhang. Preble: Efficient distributed prompt scheduling for LLM serving. InThe Thirteenth International Conference on Learning Representations, 16 ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025
2025
-
[58]
Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving
Foteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic. Déjàvu: Kv-cache stream- ing for fast, fault-tolerant generative LLM serving. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenRe- view.net, 2024
2024
-
[59]
QUEST: query-aware sparsity for efficient long-context LLM inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. QUEST: query-aware sparsity for efficient long-context LLM inference. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenRe- view.net, 2024
2024
-
[60]
Gemma 3 technical report.CoRR, abs/2503.19786, 2025
Gemma Team. Gemma 3 technical report.CoRR, abs/2503.19786, 2025
Pith/arXiv arXiv 2025
-
[61]
The llama 3 herd of models.CoRR, abs/2407.21783, 2024
Llama Team. The llama 3 herd of models.CoRR, abs/2407.21783, 2024
Pith/arXiv arXiv 2024
-
[62]
Qwen3 technical report.CoRR, abs/2505.09388, 2025
Qwen Team. Qwen3 technical report.CoRR, abs/2505.09388, 2025
Pith/arXiv arXiv 2025
-
[63]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017
2017
-
[64]
KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025
Jiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen, Wenyuan Yu, and Haibo Chen. KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025
2025
-
[65]
Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yong- wei Wu, and Congfeng Jiang. From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation.CoRR, abs/2601.12904, 2026
arXiv 2026
-
[66]
Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method
Yiming Wang, Zhuosheng Zhang, and Rui Wang. Element-aware summarization with large language models: Expert-aligned evaluation and chain-of- thought method. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 8640–8665. Association for Computa- ...
2023
-
[67]
Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation
Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen, Mingcong Han, Jinyu Gu, and Haibo Chen. Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation. InProceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13-16, 2025, pages 996–10...
2025
-
[68]
LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. LoongServe: Efficiently serv- ing long-context large language models with elastic se- quence parallelism. InProceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP 2024, Austin, TX, USA, November 4-6, 2024, pages 640–
2024
-
[69]
DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads
Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025
2025
-
[70]
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. InThe Twelfth Interna- tional Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024
2024
-
[71]
CacheBlend: Fast large language model serving for RAG with cached knowledge fusion
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yi- hua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. CacheBlend: Fast large language model serving for RAG with cached knowledge fusion. In Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Nether- lands, 30 March 2025 - 3 April 2025, pages 94–...
2025
-
[72]
FlashInfer: Efficient and customizable attention engine for LLM inference serving
Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, and Luis Ceze. FlashInfer: Efficient and customizable attention engine for LLM inference serving. 2025
2025
-
[73]
Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X
Chong You, Kan Wu, Zhipeng Jia, Lin Chen, Srinadh Bhojanapalli, Jiaxian Guo, Utku Evci, Jan Wassenberg, Praneeth Netrapalli, Jeremiah J. Willcock, Suvinay Sub- ramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X. Yu, Prateek Jain, David E. Culler, Henry M. Levy, and Sanjiv Kumar. Spark transformer: Reactivat- ing sparsity in FFN and attention.CoRR...
arXiv 2025
-
[74]
Orca: A distributed 17 serving system for transformer-based generative mod- els
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soo- jeong Kim, and Byung-Gon Chun. Orca: A distributed 17 serving system for transformer-based generative mod- els. In16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2022, Carlsbad, CA, USA, July 11-13, 2022, pages 521–538. USENIX Associa- tion, 2022
2022
-
[75]
Stateful large language model serving with Pensieve
Lingfan Yu, Jinkun Lin, and Jinyang Li. Stateful large language model serving with Pensieve. InProceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 144–158. ACM, 2025
2025
-
[76]
Yifan Yu, Yu Gan, Nikhil Sarda, Lillian Tsai, Jiaming Shen, Yanqi Zhou, Arvind Krishnamurthy, Fan Lai, Hank Levy, and David E. Culler. IC-Cache: Efficient large language model serving via in-context caching. In Proceedings of the ACM SIGOPS 31st Symposium on Op- erating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13...
2025
-
[77]
Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. Na- tive sparse attention: Hardware-aligned and natively trainable sparse attention. InProceedings of the 63rd Annual Meeting of the Association for Computation...
2025
-
[78]
Rethinking database high availability with RDMA networks.Proc
Erfan Zamanian, Xiangyao Yu, Michael Stonebraker, and Tim Kraska. Rethinking database high availability with RDMA networks.Proc. VLDB Endow., 12(11):1637– 1650, 2019
2019
-
[79]
GPU checkpoint/restore made fast and lightweight
Shaoxun Zeng, Tingxu Ren, Jiwu Shu, and Youyou Lu. GPU checkpoint/restore made fast and lightweight. In 24th USENIX Conference on File and Storage Technologies, FAST 2026, Santa Clara, CA, USA, February 24-26, 2026, pages 239–254. USENIX Association, 2026
2026
-
[80]
Jenga: Effective memory manage- ment for serving LLM with heterogeneity
Chen Zhang, Kuntai Du, Shu Liu, Woosuk Kwon, Xi- angxi Mo, Yufeng Wang, Xiaoxuan Liu, Kaichao You, Zhuohan Li, Mingsheng Long, Jidong Zhai, Joseph Gon- zalez, and Ion Stoica. Jenga: Effective memory manage- ment for serving LLM with heterogeneity. InProceed- ings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel Worl...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.