REVIEW 4 major objections 6 minor 1 cited by
Multi-agent LLM systems waste work replaying shared context as text; treating KV cache as transferable state cuts that cost sharply.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 07:50 UTC pith:5XZWTWU7
load-bearing objection Solid systems abstraction for distributed KV reuse in multi-agent workflows; headline multipliers are modeled aggregate work, not concurrent wall-clock. the 4 major comments →
[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that elevating KV cache to a workflow-level state object—with explicit materialization, transfer, fork, restricted composition, and eviction operators, plus a transfer-versus-recompute cost model—lets multi-agent LLM systems reuse long shared prefixes without repeated prefill, and that KV transmission beats recomputation on moderate-to-high bandwidth networks, producing large reductions in TTFT, multi-agent compute cost, peak KV memory, and framework overhead relative to text-centric and local-prefix baselines.
What carries the argument
Stateful operator abstraction (stateflow): each operator carries input/output data and KV state plus a state policy; the KV state object bundles model identity, configuration, block tensors, positional metadata, lineage, and placement so the runtime can safely transfer, fork (copy-on-write), resume, or fall back to recompute.
Load-bearing premise
The large reported gains rest on an analytical aggregate-cost model fed by single-prompt microbenchmarks and fixed network constants, not on measured concurrent multi-agent wall-clock runs with natural prompts and long outputs.
What would settle it
Run the same shared-prefix Tree-of-Thought or collaborative-RAG workload end-to-end on a multi-node cluster with live concurrent agents, natural prompts, and longer generations; if wall-clock TTFT and total latency show little or no gain over text replay once scheduling contention and real traffic appear, the central efficiency claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. AAFLOW+ extends the AAFLOW operator model from dataflow to stateflow by treating the LLM KV cache as a first-class distributed object with operators for materialization, transfer, fork, restricted merge, and eviction. The paper defines a structured KV state object with compatibility and lineage metadata, a hybrid data/state execution graph, and a transfer-vs-recompute cost model. Evaluation on Mistral-7B and Llama-3-8B (HF, vLLM, SGLang backends) reports large gains—up to 50.2× TTFT, 7.63× multi-agent compute cost at 16 agents, 1.72–6.10× peak KV memory, and >7.74× throughput—via an analytical aggregate-cost model parameterized by empirical microbenchmarks (prefill/decode times, measured KV bytes, bandwidth/latency). A bandwidth sweep shows transfer dominating recomputation above ~25 Gbps for the tested contexts. The central thesis is that workflow-level KV-state sharing should replace text-centric agent communication for shared-prefix multi-agent workloads.
Significance. If the qualitative result holds under real concurrent multi-agent execution, the contribution is significant for distributed LLM systems: elevating KV cache from a local serving artifact to a schedulable workflow object is a natural and useful systems abstraction, complementary to vLLM/SGLang-style local reuse. Strengths include a coherent transfer-vs-recompute cost model (Eqs. 7, 18–24, 27, 30), explicit correctness constraints (model/positional/lineage compatibility; restricted merge), a concrete bandwidth scheduling rule from Experiment 3, detailed operator algorithms in the appendix, and public artifacts. The paper is honest in §6.5 and §8 that total latency is modeled aggregate work, Y=64, and prompts are synthetic shared-prefix DAGs. Those qualifications, however, currently limit how far the headline multipliers can be taken as systems evidence.
major comments (4)
- [Abstract; §6.5; Tables 2–5] Abstract and §6 Experiments 1–5 present headline multipliers (50.2× TTFT, 7.63× multi-agent compute, >7.74× throughput) as system results, but §6.5 states that total_latency_sec is modeled aggregate work (agent×branch×prompt multiplication of microbenchmarks), not measured multi-agent wall-clock under concurrency. For a systems paper whose central claim is end-to-end efficiency of distributed KV orchestration, this gap is load-bearing: either report real concurrent multi-agent wall-clock/memory under transfer contention, or reframe all abstract/intro claims as model-predicted savings with explicit caveats matching §6.5.
- [§6.3–6.4; §8; Abstract] The evaluation regime is fixed to synthetic shared-prefix DAGs with Y=64 generated tokens per branch (§6.3–6.4; Limitations §8). Prefill dominance is then almost guaranteed, so the large TTFT and multi-agent factors are regime-specific. The abstract does not qualify this; longer decode-heavy branches or heterogeneous (non-shared) prefixes would shrink the claimed gains even if the transfer inequality remains valid. Please either add decode-length and prefix-overlap sensitivity experiments, or tightly scope every quantitative claim to short-output shared-prefix workflows.
- [§5.4–5.5; Appendix F.3; Tables 8–12] Appendix F.3 notes that vLLM and SGLang do not expose stable public KV export/import in this evaluation, so cross-node transfer is not fully exercised through production serving APIs for those backends. Combined with modeled baseline adapters (DistServe-style, KVCOMM, dense prefill), it is unclear how much of Experiments 1–2 and 5 is end-to-end runtime versus cost-model extrapolation over HF microbenchmarks. Clarify which paths actually move real KV tensors over UCX/RDMA between agents, and mark simulated vs measured rows in Tables 2–5 / 8–12.
- [§4.4; §6.9; §8; Algorithm 6] Memory claims (1.72–6.10×; Experiment 4 / Table 4) rest on fork-time block sharing under unconstrained memory; §8 states eviction and memory-constrained scheduling are only architectural objectives and not evaluated. The eviction score (Algorithm 6; α, β, γ) is therefore unvalidated. Either evaluate under GPU memory pressure with the proposed policy, or present memory results strictly as fork-sharing savings without implying a complete distributed KV memory manager.
minor comments (6)
- [Abstract] Abstract phrasing “making sure KV-state sharing greatly increases efficiency” is awkward; replace with a precise claim about when transfer beats recompute.
- [Front matter] PVLDB reference block still has placeholder year/volume/doi (2020 / XXX-XXX / XX.XX/XXX.XX).
- [References] SGLang is cited twice as [47] and [48] with identical bibliographic entries; consolidate.
- [§2–§3] Equation formatting is uneven (e.g., T_prefill, Op_kv_transfer) and some multi-line cost equations are hard to parse; normalize notation for S_in/S_out and Ω_state vs Ω_text.
- [§6 figures] Figure 7–10 captions are informative but axis units and which backend/model each curve uses should be stated in every caption, not only in nearby tables.
- [§7] Related work could briefly contrast with Mooncake, MemServe, and LMCache on whether any already expose cross-request KV as a transferable object, to sharpen novelty relative to disaggregated serving.
Circularity Check
No significant circularity: speedups are computed from empirical microbenchmarks plugged into an explicit, non-tautological transfer-vs-recompute cost model.
full rationale
The paper's central performance claims (TTFT, multi-agent aggregate cost, memory, throughput) are obtained by measuring prefill/decode times and KV byte sizes on real backends (HF/vLLM/SGLang with Mistral-7B and Llama-3-8B), then substituting those measured quantities into the open inequalities of Eqs. 7 and 19–24 (T_transfer + T_resume + Ω_state ≟ T_prefill + Ω_text, and the corresponding k-branch savings ΔT). The model is transparent that total_latency_sec is aggregate modeled work, not wall-clock, and the same model is applied uniformly to all baselines; nothing is forced by construction or by renaming a fit. Self-citation of AAFLOW supplies only the prior dataflow operator substrate that is being extended; the KV-state operators, compatibility predicates, and measured speedups do not reduce to that citation. No uniqueness theorem, ansatz smuggled via self-citation, or self-definitional loop appears. The derivation chain is therefore self-contained against the stated empirical inputs and external baselines.
Axiom & Free-Parameter Ledger
free parameters (4)
- bandwidth_bytes_per_sec (and sweep 10–400 Gbps)
- network_latency_sec / resume_overhead_sec / omega_state_sec / omega_text_sec
- output length Y = 64 tokens per branch
- eviction score weights α, β, γ
axioms (5)
- domain assumption Prefill cost dominates TTFT for long contexts and is approximately linear in L; decode is comparatively cheap for short Y.
- domain assumption Two KV states are safely reusable only under model/tokenizer identity and positional compatibility; otherwise fall back to text recompute.
- domain assumption Restricted merge is limited to non-overlapping sequential concatenation or text-level reduction; arbitrary tensor blending of divergent attention states is invalid under RoPE/positional encodings.
- ad hoc to paper Aggregate multi-agent cost can be extrapolated by multiplying per-prompt microbenchmarks by agent×branch×prompt counts (analytical model).
- domain assumption Zero-copy Arrow metadata + UCX/RDMA tensor transfer accurately models inter-node KV movement cost.
invented entities (3)
-
Stateful operator Op^s_i = (I, O, S_in, S_out, f, P, σ)
no independent evidence
-
KV state object S_KV = (M, Θ, B, Π, Λ, Γ) with block-level lineage
no independent evidence
-
Stateflow execution graph G_s = (V, E_d, E_s)
no independent evidence
read the original abstract
Multi-agent LLM systems increasingly integrate retrieval, planning, and reasoning, but remain fundamentally text-centric, requiring agents to repeatedly recompute shared context through expensive prefill. Although single-request inference is known to be accelerated by KV-cache management, it is usually restricted to local serving scopes. We introduce AAFLOW+, a stateful extension of agentic workflow operators that makes KV cache a first-class distributed systems object. AAFLOW+ builds processes into communication-aware graphs that concurrently optimize data, prompts, and reusable model state. It also provides operators for KV materialization, transfer, fork, composition, and eviction. Its runtime enables zero-copy, transfer-aware execution, allowing agents to reuse long context without recomputation. AAFLOW+ reduces TTFT by up to 50.2x, achieves up to 7.63x reduced multi-agent compute cost at 16-agent scale, reduces KV memory by 1.72-6.10x, and increases throughput by more than 7.74x, based on an analytical cost model parameterized by empirical hardware microbenchmarks. The results demonstrate that KV transmission outperforms recomputation on networks with moderate to high bandwidth, making sure KV-state sharing greatly increases efficiency in multi-agent LLM systems by replacing text passing.
Figures
Forward citations
Cited by 1 Pith paper
-
Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
Source-omitted event KV rows can carry compact upstream state (semantic materialization); deliberate answer-free carriers raise recovery from 6% to 51% on Qwen3-8B, while passive natural mentions do not.
Reference graph
Works this paper leans on
-
[1]
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee. 2023. Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills.arXiv preprint arXiv:2308.16369 (2023)
Pith/arXiv arXiv 2023
-
[2]
Apache Arrow Project. 2025. Apache Arrow. Project website. https://arrow. apache.org/ Language-independent columnar memory format with zero-copy reads
2025
-
[3]
Apache Software Foundation. 2016. Apache Arrow: A cross-language develop- ment platform for in-memory data. https://arrow.apache.org
2016
-
[4]
Katz, Ben Clifford, Rohan Kumar, Lukasz Lacinski, Ryan Chard, Justin M
Yadu Babuji, Anna Woodard, Zhuozhao Li, Daniel S. Katz, Ben Clifford, Rohan Kumar, Lukasz Lacinski, Ryan Chard, Justin M. Wozniak, Ian Foster, Michael Wilde, and Kyle Chard. 2019. Parsl: Pervasive Parallel Programming in Python. InProceedings of the 28th International Symposium on High-Performance Parallel and Distributed Computing(Phoenix, AZ, USA)(HPDC ...
doi:10.1145/33 2019
-
[5]
Paris Carbone, Asterios Katsifodimos, Stephan Ewen, Volker Markl, Seif Haridi, and Kostas Tzoumas. 2015. Apache flink: Stream and batch processing in a single engine.The Bulletin of the Technical Committee on Data Engineering38, 4 (2015). https://asterios.katsifodimos.com/assets/publications/flink-deb.pdf
2015
-
[6]
Lisandro Dalcin, Rodrigo Paz, and Mario Storti. 2005. MPI for Python.J. Parallel and Distrib. Comput.65, 9 (1 Sept. 2005), 1108–1115. https://doi.org/10.1016/j.jp dc.2005.03.010
doi:10.1016/j.jp 2005
-
[7]
Emily Davis. 2024. Building Custom AI Workflows Using LangChain Tools. ThinkTide Global Research Journal5, 4 (2024), 54–62. https://thinktidejournal.c om/index.php/TGRJ/article/view/53/63
2024
-
[8]
Jeffrey Dean and Sanjay Ghemawat. 2008. MapReduce: simplified data processing on large clusters.Commun. ACM51, 1 (Jan. 2008), 107–113. https://doi.org/10.1 145/1327452.1327492
arXiv 2008
-
[9]
Maechling, Rajiv Mayani, Weiwei Chen, Rafael Ferreira da Silva, Miron Livny, and Kent Wenger
Ewa Deelman, Karan Vahi, Gideon Juve, Mats Rynge, Scott Callaghan, Philip J. Maechling, Rajiv Mayani, Weiwei Chen, Rafael Ferreira da Silva, Miron Livny, and Kent Wenger. 2015. Pegasus, a workflow management system for science automation.Future Generation Computer Systems46 (2015), 17–35. https: //doi.org/10.1016/j.future.2014.10.008
-
[10]
2024.LangGraph: Stateful Multi-Agent Workflows
LangChain Developer. 2024.LangGraph: Stateful Multi-Agent Workflows. Techni- cal Report. LangChain Inc. https://blog.langchain.com/langgraph-multi-agent- workflows
2024
-
[11]
Yuanshuang Fu, Dan Liu, Bonan Zhang, Zhuotong Jiang, Haibo Mei, and Jiajin Guan. 2025. Cue RAG: Dynamic multi-output cue memory under H framework for retrieval-augmented generation.Neurocomputing639 (2025), 130235. https: //doi.org/10.1016/j.neucom.2025.130235
-
[12]
Yingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu, and Jiang Liu. 2026. RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse. arXiv preprint arXiv:2603.13289(2026)
arXiv 2026
-
[13]
Bernal J Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024. Hipporag: Neurobiologically inspired long-term memory for large language models.Advances in neural information processing systems37 (2024), 59532– 59569
2024
-
[14]
Bodun Hu, Jiamin Li, Le Xu, Myungjin Lee, Akshay Jajoo, Geon-Woo Kim, Hong Xu, and Aditya Akella. 2024. Blockllm: Multi-tenant finer-grained serving for large language models.arXiv preprint arXiv:2404.18322(2024)
Pith/arXiv arXiv 2024
-
[15]
Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan. 2024. Mem- Serve: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.arXiv preprint arXiv:2406.17565(2024). https://arxiv.org/pdf/2406.17565
Pith/arXiv arXiv 2024
-
[16]
Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav San- thanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2023. DSPy: Com- piling Declarative Language Model Calls into Self-Improving Pipelines.arXiv preprint arXiv:2310.03714(2023). https://arxiv.or...
Pith/arXiv arXiv 2023
-
[17]
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics7 (2019), 453–466
2019
-
[18]
Woosuk Kwon et al. 2023. vLLM: Easy, Fast, and Cheap LLM Serving. https: //github.com/vllm-project/vllm
2023
-
[19]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principles(Koblenz, Germany)(SOSP ’23). Association for Computing Machinery, New York, N...
-
[20]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. InProceedings of the 34th Inter- national Conference on Neural Information Processing Systems(V...
-
[21]
Yuhan Liu, Jiayi Yao, Yihua Cheng, Yuwei An, Xiaokun Chen, Shaoting Feng, et al
-
[22]
arXiv preprint arXiv:2510.09665(2025)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference. arXiv preprint arXiv:2510.09665(2025). https://arxiv.org/pdf/2510.09665
arXiv 2025
-
[23]
LLMS3. 2026. When AI Memory Became an Architecture: KV-Cache Persistence, MCP, and the Night S3 Got Its Memory Tier. Project website. https://llms3. com/blog/when-ai-memory-became-an-architecture-may-2026 KV-Cache Persistence, MCP, and the Night S3 Got Its Memory Tier
2026
-
[24]
A. Merzky, M. Turilli, M. Titov, A. Al-Saadi, and S. Jha. 2022. Design and Perfor- mance Characterization of RADICAL-Pilot on Leadership-Class Platforms.IEEE Transactions on Parallel and amp; Distributed Systems33, 04 (apr 2022), 818–829. https://doi.org/10.1109/TPDS.2021.3105994
-
[25]
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I Jordan, et al. 2018. Ray: A distributed framework for emerging {AI} applications. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 561–577. https://www.usenix.org/system/files/os...
2018
-
[26]
Niranda Perera, Arup Kumar Sarker, Kaiying Shan, Alex Fetea, Supun Kambu- rugamuve, Thejaka Amila Kanewala, Chathura Widanage, Mills Staylor, Tianle Zhong, Vibhatha Abeykoon, Gregor von Laszewski, and Geoffrey Fox. 2024. Supercharging distributed computing environments for high-performance data engineering.Frontiers in High Performance ComputingVolume 2 -...
-
[27]
Niranda Perera, Arup Kumar Sarker, Mills Staylor, Gregor von Laszewski, Kaiy- ing Shan, Supun Kamburugamuve, Chathura Widanage, Vibhatha Abeykoon, Thejaka Amila Kanewela, and Geoffrey Fox. 2023. In-depth analysis on parallel processing patterns for high-performance Dataframes.Future Generation Com- puter Systems149 (2023), 250–264. https://doi.org/10.1016...
-
[28]
Maximilian Petersohn, Stephen Macke, Doris Xin, William Ma, J. K. Wittenauer, Stephen Hoyer, Ryan Marcus, Matei Zaharia, and Benjamin Recht. 2020. Towards Scalable Dataframe Systems.Proceedings of the VLDB Endowment (PVLDB)13, 12 (2020), 2033–2046. https://doi.org/10.14778/3407790.3407807
-
[29]
Hongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao, Defu Lian, Zhicheng Dou, and Tiejun Huang. 2025. MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation. InProceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(WWW ’25). Association for Computing Machinery, New York, NY, USA, 2366–2377. https://doi.or...
arXiv 2025
-
[30]
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Heyi Tang, Feng Ren, Teng Ma, Shangming Cai, Yineng Zhang, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.ACM Trans. Storage(Nov. 2025). https://doi.org/10.1145/3773 772 Just Accepted
-
[31]
Matthew Rocklin. 2015. Dask: Parallel computation with blocked algorithms and task scheduling. InProceedings of the 14th python in science conference, Vol. 130. Citeseer, 136. https://proceedings.scipy.org/articles/Majora-7b98e3ed-013.pdf
2015
-
[32]
Arup Kumar Sarker, Aymen Alsaadi, Alexander James Halpern, Prabhath Tan- gella, Mikhail Titov, Niranda Perera, Mills Staylor, Gregor von Laszewski, Shantenu Jha, and Geoffrey Fox. 2025. Deep RC: A Scalable Data Engineering and Deep Learning Pipeline. InJob Scheduling Strategies for Parallel Process- ing: 28th International Workshop, JSSPP 2025, Milan, Ita...
-
[33]
Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera, Mills Staylor, Gregor von Laszewski, Matteo Turilli, Ozgur Ozan Kilic, Mikhail Titov, Andre Merzky, Shantenu Jha, et al. 2024. Design and implementation of an analysis pipeline for heterogeneous data.arXiv preprint arXiv:2403.15721(2024)
Pith/arXiv arXiv 2024
-
[34]
Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera, Mills Staylor, Gregor von Laszewski, Matteo Turilli, Ozgur Ozan Kilic, Mikhail Titov, Andre Merzky, Shantenu Jha, et al. 2024. Radical-Cylon: A Heterogeneous Data Pipeline for Scientific Computing. InJob Scheduling Strategies for Parallel Processing. Springer Nature Switzerland, 84–102. https://doi.org/10....
-
[35]
Arup Kumar Sarker, Mills Staylor, Aymen Alsaadi, Gregor von Laszewski, Shantenu Jha, and Geoffrey Fox. 2026. AAFLOW: Scalable Patterns for Agentic AI Workflows.arXiv preprint arXiv:2605.02162(2026). Under Submission to SC2026
Pith/arXiv arXiv 2026
-
[36]
Pavel Shamis, Manjunath Gorentla Venkata, M. Graham Lopez, Matthew B. Baker, Oscar Hernandez, Yossi Itigin, Mike Dubman, Gilad Shainer, Richard L. Graham, Liran Liss, Yiftah Shahar, Sreeram Potluri, Davide Rossetti, Donald Becker, Duncan Poole, Christopher Lamb, Sameer Kumar, Craig Stunkel, George Bosilca, and Aurelien Bouteiller. 2015. UCX: An Open Sourc...
-
[37]
Kaiying Shan, Niranda Perera, Damitha Lenadora, Tianle Zhong, Arup Ku- mar Sarker, Supun Kamburugamuve, Thejaka Amila Kanewela, Chathura Widan- age, and Geoffrey Fox. 2022. Hybrid Cloud and HPC Approach to High- Performance Dataframes. In2022 IEEE International Conference on Big Data (Big Data). 2728–2736. https://doi.org/10.1109/BigData55660.2022.10020958
-
[38]
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, and Ce Zhang. 2023. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. InProceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research), Andreas Krau...
2023
-
[39]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: language agents with verbal reinforcement learning. InProceedings of the 37th International Conference on Neural Information Process- ing Systems(New Orleans, LA, USA)(NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, Article 377, 19 pages. https://dl.a...
-
[40]
Mills Staylor, Arup Kumar Sarker, Gregor von Laszewski, Geoffrey Fox, Yue Cheng, and Judy Fox. 2026. Combining Serverless and High-Performance Com- puting Paradigms to support ML Data-Intensive Applications.Frontiers in High Performance Computing(2026)
2026
-
[41]
Chathura Widanage, Niranda Perera, Vibhatha Abeykoon, Supun Kamburuga- muve, Thejaka Amila Kanewala, Hasara Maithree, Pulasthi Wickramasinghe, Ahmet Uyar, Gurhan Gunduz, and Geoffrey Fox. 2020. High performance data engineering everywhere. In2020 IEEE International Conference on Smart Data Services (SMDS). IEEE, 122–132. https://doi.org/10.1109/SMDS49396....
-
[42]
White, Doug Burger, and Chi Wang
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang (Eric) Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Ahmed Awadallah, Ryen W. White, Doug Burger, and Chi Wang. 2024. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. InCOLM 2024. https://www.mi crosoft.com/en-us/research/publication/autogen-enabling-next-gen-llm- a...
2024
-
[43]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InInternational Conference on Learning Representations (ICLR). https: //openreview.net/pdf?id=WE_vluYUL-X
2023
-
[44]
Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Zhijian Liu, Jianyi Zhang, Danyang Zhuo, and Yiran Chen
-
[45]
KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
-
[46]
Lu Ye, Ze Tao, Yong Huang, and Yang Li. 2024. ChunkAttention: Efficient Self- Attention with Prefix-Aware KV Cache and Two-Phase Partition. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). 11608–11620. https://doi.org/10.18653/v1/2024.acl-long.623
-
[47]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung- Gon Chun. 2022. Orca: A distributed serving system for {Transformer-Based} generative models. In16th USENIX symposium on operating systems design and implementation (OSDI 22). 521–538
2022
-
[48]
Matei Zaharia, Reynold S. Xin, Patrick Wendell, Tathagata Das, Michael Armbrust, Ankur Dave, Xiangrui Meng, Josh Rosen, Shivaram Venkataraman, Michael J. Franklin, Ali Ghodsi, Joseph Gonzalez, Scott Shenker, and Ion Stoica. 2016. Apache Spark: a unified engine for big data processing.Commun. ACM59, 11 (Oct. 2016), 56–65. https://doi.org/10.1145/2934664
doi:10.1145/2934664 2016
-
[50]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: efficient execution of struc- tured language model programs. InProceedings of the 38th International Con- ference on Neural Information Processing Systems(Vancouver, B...
-
[51]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI). https://www.usenix.o rg/conference/osdi24/presentation/zhong-yinmin
2024
-
[52]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 193–210. APPENDIX A EXTENDED SYSTEM DESIGN This appendix expands the de...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.