REVIEW 3 major objections 4 minor 251 references
State management is a coupled control loop, not a storage problem.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:46 UTC pith:FTFDLKO4
load-bearing objection A genuinely useful cross-domain survey of state management whose central 'coupled control loop' claim is under-built formally but survives on the strength of its comparative examples. the 3 major comments →
Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central discovery is that state management is a propagation problem, not a pipeline: state transitions follow S_{t+1}=V(E(A(S_t;W_t,Γ_t);H);U_t), where A, E, V are access, execution, and evolution ops, and the paper asserts these are nonseparable (∂E/∂A, ∂V/∂E, ∂A/∂V are all nonzero). A local decision in any stage returns as a constraint in the others—access skew destroys locality, execution shortcuts create movement or quality debt, aggressive updates become future contention. The recurring bottleneck in real systems, it argues, is not missing algorithms but missing contracts: no agreed specification of which state object is governed, which actions are legal, which service bound
What carries the argument
The central objects are (1) the propagation structure S_{t+1}=V(E(A(S_t;W_t,Γ_t);H);U_t), which formalizes state management as a coupled loop with nonseparable access (A), execution (E), and evolution (V) stages, and (2) the reusable five-field comparison tuple Q=(σ,κ,χ,β,γ) — state object, control surface, coupling path, evaluation boundary, unresolved contract — which normalizes papers across domains so mechanisms can be compared and transferred. The propagation structure carries the paper's core argument (local decisions are never local), while the tuple is the analytical scaffold that makes cross-domain comparison possible.
Load-bearing premise
The central claim rests on the assumption that access, execution, and evolution are never separable in real systems; if important workloads keep these stages independent, the propagation view overstates coupling.
What would settle it
Run a controlled experiment on a stateful system with short-lived state (e.g., a serverless function platform): hold workload and update policy fixed, change only the access-scheduling policy, and measure whether execution cost and update/retention cost change. If a range of workloads shows zero cross-stage effect, the nonseparability claim fails.
If this is right
- Stateful runtimes should expose typed state catalogs, lifecycle stages (admit, place, mutate, expose/transfer, compact, evict/reclaim), and explicit contracts; without them, local gains create lifecycle debt.
- Evaluations must move from one-shot benchmarks to multi-phase disturbance traces (warmup, burst, drift, reconfiguration, recovery) to reveal debt that stationary runs hide.
- Mechanisms transfer across streaming, serving, retrieval, and retention only when their five-field contracts align; a policy that works in one domain may fail when state identity, ownership, or disturbance horizon differs.
- The recurring bottleneck in practice is missing contracts, not missing algorithms—so design reviews should focus on legal actions, protected boundaries, and charged costs.
Where Pith is reading between the lines
- A quantitative version of the nonseparability claim could be tested by estimating cross-derivatives of an end-to-end cost function from system measurements; the paper asserts the derivatives are nonzero but does not measure them.
- The five-field tuple could double as a reviewer checklist for future systems papers, forcing authors to state their state object and evaluation boundary explicitly.
- If the framework is right, an architectural opportunity follows: a shared middleware layer that unifies observability and actuation over state across streaming, serving, and retrieval—a direction the paper names but does not build.
- The framework implies that 'state debt' accounting (e.g., charging restoration or rebuild cost to the same SLO boundary as the immediate gain) could be made a first-class runtime metric; this is not yet implemented in any surveyed system.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys parallel and distributed systems literature on shared state, organizing it around three coupled dimensions—state access and scheduling, state-aware execution, and state evolution and reuse. It proposes a five-field comparison tuple (state object, control surface, coupling path, evaluation boundary, unresolved contract) and applies it across streaming, LLM serving, retrieval, and continual-learning systems. The central claim, stated in §§2.5.2 and 7, is that efficient state management is best understood as one coupled control loop over these three dimensions. The paper also contributes a contract-oriented blueprint for stateful runtimes and a disturbance-aware evaluation agenda.
Significance. The survey offers a useful vocabulary and a broad cross-domain synthesis. If accepted, it would help unify several literatures and focus evaluation on lifecycle debt. The paper is honest about evidence classes and marks unresolved contracts, which is a strength. However, the central coupling claim is not established: the only formal support is a composition equation whose nonseparability conditions are tautological under standard function-composition semantics. The qualitative running examples are plausible but do not constitute a systematic test. The contribution is therefore best viewed as a well-organized research agenda and taxonomy rather than an established theorem. With a precise definition of separability and either empirical evidence or a restated, more modest claim, the paper would be valuable.
major comments (3)
- [§2.5.2] The claim ∂E/∂A≠0, ∂V/∂E≠0, ∂A/∂V≠0 is presented as the core evidence for nonseparability. Under the displayed composition S_{t+1}=V(E(A(S_t;W_t,Γ_t);H);U_t), these inequalities follow from A's output feeding E and E's output feeding V, and ∂A/∂V from the next time step's A reading V's output. They are entailed by the notation and cannot distinguish substantive cross-stage feedback from mere sequential dataflow. For the conclusion 'best understood as one coupled control loop' to carry weight, the paper needs a precise definition of separability (e.g., whether the joint optimum equals composition of per-stage optima, or whether stage policies are invariant under other stages' perturbations) and evidence that real systems violate it. As written, the formal anchor is tautological and the central claim rests on selected examples.
- [§2.1] The inclusion criterion 'Papers enter the analytical core when they externalize a governable runtime seam over state' selects for systems that fit the control-loop narrative. This makes the conclusion that state management is a control problem partly circular: systems that do not expose such a seam are treated as contextual, so the corpus is biased toward the thesis. The paper should either demonstrate that the excluded systems are rare or irrelevant, or reframe the conclusion as 'for systems with exposed state-control seams, the coupled-loop view is useful' rather than a general claim about state management.
- [§3.4] The quantitative propagation trace uses 'representative numbers drawn from the vLLM and Sarathi-Serve literature' (see the paragraph beginning 'Consider a multi-tenant LLM serving cluster'). This is a scenario with hypothetical numbers, not a measurement or a test of the nonseparability claim. The trace illustrates how admission could propagate to restoration debt, but it does not establish that the three stages are generally coupled or that the coupling is control-relevant in the surveyed systems. Either label this as a motivating example and soften the conclusion, or provide an actual empirical study measuring propagation across the three stages.
minor comments (4)
- [§2.4] The formal sketch R=(W,P,Σ,Γ) is introduced but never used in later analysis; tie it to the propagation equation in §2.5.2 or drop it for clarity.
- [§2.6.1] The term 'unresolved contract' is central but not formally defined when first introduced; consider giving a one-sentence definition at Table 3.
- [Figure 1 caption] The caption claims the stages are 'not separable pipeline stages'; given the formal issue in §2.5.2, this should be rephrased as 'we hypothesize' or accompanied by empirical evidence.
- [Appendix B.7] The policy skeleton in Table 10 includes functions like observe_signals() and account_outcome(); clarify whether this is pseudocode or a conceptual template, and how it maps to the blueprint layers in §5.3.
Circularity Check
§2.5.2's nonseparability conditions are entailed by the paper's own composition notation, making the formal anchor of the 'coupled control loop' thesis tautological; the qualitative synthesis provides independent but weaker support.
specific steps
-
self definitional
[Section 2.5.2, 'The Propagation Perspective']
"Propagation structure. Let the runtime state at logical time t be denoted S_t. Define three stage operators—access (A), execution (E), and evolution (V)—each producing the next state snapshot: S_{t+1}=V(E(A(S_t;W_t,Γ_t);H);U_t). The core claim is that these stages are not separable: ∂E/∂A≠0, ∂V/∂E≠0, ∂A/∂V≠0."
A, E, and V are defined sequentially: E takes A's output as its first argument, V takes E's output, and the next A takes V's output. Under the standard reading of partial derivatives, ∂E/∂A≠0 etc. hold for essentially any function that reads its input; they express the composition itself, not a measured or derived property of real systems. The notation therefore builds in the 'not separable' conclusion. To show nonseparability the paper would need an independent definition of separability (e.g., joint optimum equals composition of per-stage optima) or cross-partial derivatives from a concrete model; neither is given. The 'core claim' is a restatement of the defining equation, so the formal anchor is circular.
full rationale
The paper is a broad survey whose main value is the five-field tuple and the cross-domain synthesis; that synthesis is genuinely comparative and cites many external systems. The self-citations (Spacker, MorphStream, StreamFP, SAGE, KELDAR, HeterRAG, GRACE) are supporting examples, not the sole or load-bearing evidence for any claim. The inclusion criteria are explicit scope choices and do not by themselves force the conclusion. However, the central thesis's formal anchor in §2.5.2 is circular: the compound expression S_{t+1}=V(E(A(...));...) already writes the stages as a composition, so the asserted nonseparability ('∂E/∂A≠0' etc.) follows from the notation, not from observation or a model. A definition of separability independent of the composition (e.g., whether stage-wise optimal policies compose to a joint optimum) is absent. Thus the 'coupled control loop' claim is not independently derived, though the qualitative running examples give it partial empirical support. Overall: one significant formal circularity, but the survey has independent descriptive content, so score 5 rather than 8-10.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Managed state requires Lifetime, Influence, and Controllable (conditions (1)-(3) in §2.4).
- domain assumption The access/execution/evolution stages are non-separable: ∂E/∂A≠0, ∂V/∂E≠0, ∂A/∂V≠0 (§2.5.2).
- ad hoc to paper The corpus inclusion criteria (state visibility, exposed mechanism, end-to-end connection) are sufficient to support cross-domain generalization (§2.1).
invented entities (2)
-
five-field comparison tuple (state object, control surface, coupling path, evaluation boundary, unresolved contract)
no independent evidence
-
lifecycle debt / restoration debt
no independent evidence
read the original abstract
Shared state increasingly shapes both performance and failure behavior in streaming, serving, retrieval, and continual-learning systems. Existing studies, however, often isolate access control, hardware-aware execution, memory management, and long-horizon updates. The review organizes this literature around three coupled dimensions: state access and scheduling, state-aware execution, and state evolution and reuse. Across these dimensions, the literature is synthesized through a common scaffold: state object, control surface, coupling path, evaluation boundary, and unresolved contract. This comparison identifies recurring anti-patterns and informs a contract-oriented blueprint and disturbance-aware evaluation agenda. Taken together, the evidence characterizes state management as a runtime control problem.
Figures
Reference graph
Works this paper leans on
-
[1]
Daniel J. Abadi, Yanif Ahmad, Magdalena Balazinska, Ugur Cetintemel, Mitch Cherniack, Jeong-Hyon Hwang, Wolfgang Lindner, Anurag Maskey, Alex Rasin, Esther Ryvkina, Nesime Tatbul, Ying Xing, and Stan Zdonik. 2005. The Design of the Borealis Stream Processing Engine. InProceedings of the Second Biennial Conference on Innovative Data Systems Research (CIDR)...
2005
-
[2]
Daniel J. Abadi, Don Carney, Ugur Çetintemel, Mitch Cherniack, Christian Convey, Sangdon Lee, Michael Stonebraker, Nesime Tatbul, and Stan Zdonik. 2003. Aurora: a new model and architecture for data stream management.The VLDB Journal12, 2 (Aug. 2003), 120–139. doi:10.1007/s00778- 003-0095-z
doi:10.1007/s00778- 2003
-
[3]
Nair, Ilya Soloveychik, and Purushotham Kamath
Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant J. Nair, Ilya Soloveychik, and Purushotham Kamath. 2024. Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference. InProceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.), Vol. 6. 114–127. https://proceedings.mlsys.org/paper_fi...
2024
-
[4]
Saurabh Agarwal, Bodun Hu, Anyong Mao, Aditya Akella, and Shivaram Venkataraman. 2026. SYMPHONY: Enabling Compute-Memory Disaggregation in LLM Serving Systems. In23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26). USENIX Association, Renton, WA, 2027–2041. https://www.usenix.org/conference/nsdi26/presentation/agarwal
2026
-
[5]
Sameer Agarwal, Barzan Mozafari, Aurojit Panda, Henry Milner, Samuel Madden, and Ion Stoica. 2013. BlinkDB: queries with bounded errors and bounded response times on very large data. InProceedings of the 8th ACM European Conference on Computer Systems(Prague, Czech Republic) (EuroSys ’13). Association for Computing Machinery, New York, NY, USA, 29–42. doi...
arXiv 2013
-
[6]
Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee
-
[7]
Tyler Akidau, Alex Balikov, Kaya Bekiroğlu, Slava Chernyak, Josh Haberman, Reuven Lax, Sam McVeety, Daniel Mills, Paul Nordstrom, and Sam Whittle. 2013. MillWheel: fault-tolerant stream processing at internet scale.Proc. VLDB Endow.6, 11 (Aug. 2013), 1033–1044. doi:10.14778/2536222. 2536229
-
[8]
Tyler Akidau, Robert Bradshaw, Craig Chambers, Slava Chernyak, Rafael J. Fernández-Moctezuma, Reuven Lax, Sam McVeety, Daniel Mills, Frances Perry, Eric Schmidt, and Sam Whittle. 2015. The dataflow model: a practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing.Proc. VLDB Endow.8, 12 (Aug. ...
arXiv 2015
-
[9]
Waleed Ali, Siti Mariyam Shamsuddin, and Abdul Samad Ismail. 2011. A Survey of Web Caching and Prefetching A Survey of Web Caching and Prefetching.International Journal of Advances in Soft Computing and its Applications3 (03 2011)
2011
-
[10]
Rahaf Aljundi, Eugene Belilovsky, Tinne Tuytelaars, Laurent Charlin, Massimo Caccia, Min Lin, and Lucas Page-Caccia. 2019. Online Continual Learning with Maximal Interfered Retrieval. InAdvances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, Appendix 39 F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran As...
2019
-
[11]
Daiyaan Arfeen, Zhen Zhang, Xinwei Fu, Gregory Ganger, and Yida Wang. 2025. PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training. InProceedings of Machine Learning and Systems, M. Zaharia, G. Joshi, and Y. Lin (Eds.), Vol. 7. MLSys. https://proceedings.mlsys.org/ paper_files/paper/2025/file/53d3f45797970d323bd8a0d379c525aa-Paper-Conference.pdf
2025
-
[12]
Michael Armbrust, Tathagata Das, Joseph Torres, Burak Yavuz, Shixiong Zhu, Reynold Xin, Ali Ghodsi, Ion Stoica, and Matei Zaharia. 2018. Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark. InProceedings of the 2018 International Conference on Management of Data(Houston, TX, USA)(SIGMOD ’18). Association for Computing Machin...
doi:10.1145/3183713 2018
-
[13]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avi Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. InInternational Conference on Learning Representations, B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (Eds.), Vol. 2024. 9112–9141. https://proceedings.iclr.cc/paper_files/paper/202...
2024
-
[14]
Corbett, JJ Furman, Andrey Khorlin, James Larson, Jean-Michel Leon, Yawei Li, Alexander Lloyd, and Vadim Yushprakh
Jason Baker, Chris Bond, James C. Corbett, JJ Furman, Andrey Khorlin, James Larson, Jean-Michel Leon, Yawei Li, Alexander Lloyd, and Vadim Yushprakh. 2011. Megastore: Providing Scalable, Highly Available Storage for Interactive Services. InProceedings of the Conference on Innovative Data system Research (CIDR). 223–234. http://www.cidrdb.org/cidr2011/Pape...
2011
-
[15]
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego De Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, ...
2022
-
[16]
Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and SIMONE CALDERARA. 2020. Dark Experience for General Continual Learning: a Strong, Simple Baseline. InAdvances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 15920–15930. https://proceedings...
2020
-
[19]
Wei Cao, Yang Liu, Zhushi Cheng, Ning Zheng, Wei Li, Wenjie Wu, Linqiang Ouyang, Peng Wang, Yijing Wang, Ray Kuan, Zhenjun Liu, Feng Zhu, and Tong Zhang. 2020. POLARDB Meets Computational Storage: Efficiently Support Analytical Workloads in Cloud-Native Relational Database. In18th USENIX Conference on File and Storage Technologies (FAST 20). USENIX Associ...
2020
-
[20]
Paris Carbone, Gyula Fóra, Stephan Ewen, Seif Haridi, and Kostas Tzoumas. 2015. Lightweight Asynchronous Snapshots for Distributed Dataflows. (2015). doi:10.48550/arXiv.1506.08603
-
[21]
Paris Carbone, Asterios Katsifodimos, Stephan Ewen, Volker Markl, Seif Haridi, and Kostas Tzoumas. 2015. Apache flink: Stream and batch processing in a single engine.The Bulletin of the Technical Committee on Data Engineering38, 4 (2015)
2015
-
[22]
Raul Castro Fernandez, Matteo Migliavacca, Evangelia Kalyvianaki, and Peter Pietzuch. 2013. Integrating scale out and fault tolerance in stream processing using operator state management. InProceedings of the 2013 ACM SIGMOD International Conference on Management of Data(New York, New York, USA)(SIGMOD ’13). Association for Computing Machinery, New York, ...
arXiv 2013
-
[23]
Henry, Robert Bradshaw, and Nathan Weizenbaum
Craig Chambers, Ashish Raniwala, Frances Perry, Stephen Adams, Robert R. Henry, Robert Bradshaw, and Nathan Weizenbaum. 2010. FlumeJava: easy, efficient data-parallel pipelines. InProceedings of the 31st ACM SIGPLAN Conference on Programming Language Design and Implementation (Toronto, Ontario, Canada)(PLDI ’10). Association for Computing Machinery, New Y...
arXiv 2010
-
[24]
Platt, James F
Badrish Chandramouli, Jonathan Goldstein, Mike Barnett, Robert DeLine, Danyel Fisher, John C. Platt, James F. Terwilliger, and John Wernsing
-
[25]
Mani Chandy and Leslie Lamport
K. Mani Chandy and Leslie Lamport. 1985. Distributed snapshots: determining global states of distributed systems.ACM Trans. Comput. Syst.3, 1 (Feb. 1985), 63–75. doi:10.1145/214451.214456
arXiv 1985
-
[27]
Fay Chang, Jeffrey Dean, Sanjay Ghemawat, Wilson C. Hsieh, Deborah A. Wallach, Mike Burrows, Tushar Chandra, Andrew Fikes, and Robert E. Gruber. 2008. Bigtable: A Distributed Storage System for Structured Data.ACM Trans. Comput. Syst.26, 2, Article 4 (June 2008), 26 pages. doi:10.1145/1365815.1365816
arXiv 2008
-
[28]
Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. 2019. Efficient Lifelong Learning with A-GEM. InInternational Conference on Learning Representations. https://openreview.net/forum?id=Hkf2_sC5FX
2019
-
[29]
Cheng Chen, Chenzhe Jin, Yunan Zhang, Sasha Podolsky, Chun Wu, Szu-Po Wang, Eric Hanson, Zhou Sun, Robert Walzer, and Jianguo Wang. 2024. SingleStore-V: An Integrated Vector Database System in SingleStore.Proc. VLDB Endow.17, 12 (Aug. 2024), 3772–3785. doi:10.14778/3685800.3685805
arXiv 2024
-
[31]
Lequn Chen, Zihao Ye, Yongji Wu, Danyang Zhuo, Luis Ceze, and Arvind Krishnamurthy. 2024. Punica: Multi-Tenant LoRA Serving. InProceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.), Vol. 6. 1–13. https://proceedings.mlsys.org/paper_files/paper/ 2024/file/054de805fcceb78a201f5e9d53c85908-Paper-Conference.pdf
2024
-
[33]
Peizhuang Cong, Tong Yang, Yuchao Zhang, Wendong Wang, and Ke Xu. 2026. MICO: efficient query scheduling for multi-cloud deployed LLM inference service.Science China Information Sciences69, 3 (2026), 132102
2026
-
[34]
Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, J
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson Hsieh, Sebastian Kanthak, Eugene Kogan, Hongyi Li, Alexander Lloyd, Sergey Melnik, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak,...
2013
-
[35]
Daniel Crankshaw, Gur-Eyal Sela, Xiangxi Mo, Corey Zumar, Ion Stoica, Joseph Gonzalez, and Alexey Tumanov. 2020. InferLine: latency-aware provisioning and scaling for prediction serving pipelines. InProceedings of the 11th ACM Symposium on Cloud Computing(Virtual Event, USA) (SoCC ’20). Association for Computing Machinery, New York, NY, USA, 477–491. doi:...
arXiv 2020
-
[36]
Franklin, Joseph E
Daniel Crankshaw, Xin Wang, Giulio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017. Clipper: a low-latency online prediction serving system. InProceedings of the 14th USENIX Conference on Networked Systems Design and Implementation(Boston, MA, USA)(NSDI’17). USENIX Association, USA, 613–627
2017
-
[37]
Gianpaolo Cugola and Alessandro Margara. 2012. Processing flows of information: From data stream to complex event processing.ACM Comput. Surv.44, 3, Article 15 (June 2012), 62 pages. doi:10.1145/2187671.2187677
arXiv 2012
-
[38]
Tri Dao. 2024. Flashattention-2: Faster attention with better parallelism and work partitioning. InInternational Conference on Learning Representations, Vol. 2024. 35549–35562
2024
-
[39]
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. 2022. A Continual Learning Survey: Defying Forgetting in Classification Tasks.IEEE Transactions on Pattern Analysis and Machine Intelligence44, 7 (2022), 3366–3385. doi:10.1109/TPAMI.2021.3057446
arXiv 2022
-
[40]
Giuseppe DeCandia, Deniz Hastorun, Madan Jampani, Gunavardhan Kakulapati, Avinash Lakshman, Alex Pilchin, Swaminathan Sivasubramanian, Peter Vosshall, and Werner Vogels. 2007. Dynamo: amazon’s highly available key-value store. InProceedings of Twenty-First ACM SIGOPS Symposium on Operating Systems Principles(Stevenson, Washington, USA)(SOSP ’07). Associat...
arXiv 2007
-
[41]
David Domingo and Sudarsun Kannan. 2021. pFSCK: Accelerating File System Checking and Repair for Modern Storage. In19th USENIX Conference on File and Storage Technologies (FAST 21). USENIX Association, 113–126. https://www.usenix.org/conference/fast21/presentation/domingo
2021
-
[42]
Juechu Dong, Jonah Rosenblum, and Satish Narayanasamy. 2025. Toleo: Scaling Freshness to Tera-scale Memory Using CXL and PIM. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4(Hilton La Jolla Torrey Pines, La Jolla, CA, USA)(ASPLOS ’24). Association for Computing Machi...
arXiv 2025
-
[43]
Aleksandar Dragojević, Dushyanth Narayanan, Orion Hodson, and Miguel Castro. 2014. FaRM: fast remote memory. InProceedings of the 11th USENIX Conference on Networked Systems Design and Implementation(Seattle, WA)(NSDI’14). USENIX Association, USA, 401–414
2014
-
[44]
Kuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng, Qing Lan, Hejian Sang, Yihua Cheng, Jiayi Yao, Xiaoxuan Liu, Yifan Qiao, Ion Stoica, and Junchen Jiang. 2025. PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications. (2025). doi:10.48550/ ARXIV.2505.07203
-
[45]
Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali Annavaram. 2022. Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models. In19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). USENIX Association, Renton...
2022
-
[48]
Fei Fang, Yifan Hua, Shengze Wang, Ruilin Zhou, Yi Liu, Chen Qian, and Xiaoxue Zhang. 2026. PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving. In23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26). USENIX Association, Renton, WA, 2111–2129. https://www.usenix.or...
2026
-
[49]
Marco Federici, Davide Belli, Mart Van Baalen, Amir Jalalirad, Andrii Skliar, Bence Major, Markus Nagel, and Paul Whatmough. 2025. Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking. InProceedings of Machine Learning and Systems, M. Zaharia, G. Joshi, and Y. Lin (Eds.), Vol. 7. MLSys. https://proceedings.mlsys.org/paper_files/pape...
2025
-
[50]
Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, and Jie Wu. 2025. WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing Machinery, New York, NY, USA, 1283–1295. doi:10.1145/3695053.3730999
arXiv 2025
-
[51]
Weiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng, Haibin Lin, and Minlan Yu. 2025. Optimus: accelerating large-scale multi-modal LLM training by bubble exploitation. InProceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference(Boston, MA, USA)(USENIX ATC ’25). USENIX Association, USA, Article 10, 17 pages
2025
-
[52]
Xiang Fu, Weiping Zhang, Shiman Meng, Xin Huang, Wubiao Xu, Luanzheng Guo, and Kento Sato. 2024. AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis. InProceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis(Atlanta, GA, USA)(SC ’24). IEEE Press, Article 99, 16 ...
Pith/arXiv arXiv 2024
-
[53]
Yaosheng Fu, Evgeny Bolotin, Aamer Jaleel, Gal Dalal, Shie Mannor, Jacob Subag, Noam Korem, Michael Behar, and David Nellans. 2023. AutoScratch: ML-Optimized Cache Management for Inference-Oriented GPUs. InProceedings of Machine Learning and Systems, D. Song, M. Carbin, and T. Chen (Eds.), Vol. 5. Curan, 495–512. https://proceedings.mlsys.org/paper_files/...
2023
-
[54]
Yuqi Fu, Li Liu, Haoliang Wang, Yue Cheng, and Songqing Chen. 2022. SFS: smart OS scheduling for serverless functions. InProceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis(Dallas, Texas)(SC ’22). IEEE Press, Article 42, 16 pages
2022
-
[55]
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024. ServerlessLLM: low-latency serverless inference for large language models. InProceedings of the 18th USENIX Conference on Operating Systems Design and Implementation (Santa Clara, CA, USA)(OSDI’24). USENIX Association, USA, Article 8, 19 pages
2024
-
[56]
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024. Cost-efficient large language model serving for multi-turn conversations with CachedAttention. InProceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference(Santa Clara, CA, USA)(USENIX ATC’24). USENIX Association...
2024
-
[57]
Hongru Gao, Shuhao Zhang, Xiaofei Liao, and Hai Jin. 2026. GRACE: Alleviating Reconstruction Cost in Dynamic Graph Processing Systems. In Proceedings of the IEEE International Conference on Data Engineering
2026
-
[59]
Shiwei Gao, Qing Wang, Shaoxun Zeng, Youyou Lu, and Jiwu Shu. 2025. Weaver: Efficient Multi-LLM Serving with Attention Offloading. In2025 USENIX Annual Technical Conference (USENIX ATC 25). USENIX Association, Boston, MA, 587–595. https://www.usenix.org/conference/atc25/ presentation/gao
2025
-
[60]
Xin Gao, Sibasish Acharya, Sihui Han, Yongxiong Ren, Yanli Zhao, Liang Luo, Chucheng Wang, Pradeep Fernando, Saurabh Mishra, Siqi Yan, Yicong Du, Elzbieta Krepska, Intaik Park, Min Ni, Qunshu Zhang, and Shen Li. 2025. DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems.Proc. VLDB Endow.18, 12 (Aug. 2025), 4978–4990. doi:10.14778...
arXiv 2025
-
[62]
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2024. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. InProceedings of Machine Learning and Systems, P. Gibbons, G. Pekhimenko, and C. De Sa (Eds.), Vol. 6. 325–338. https://proceedings.mlsys.org/paper_files/paper/2024/file/a66caa1703fe34705a4368c3014c196...
2024
-
[64]
Inigo Goiri, Ricardo Bianchini, Santosh Nagarakatte, and Thu D. Nguyen. 2015. ApproxHadoop: Bringing Approximations to MapReduce Frameworks. InProceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems (Istanbul, Turkey)(ASPLOS ’15). Association for Computing Machinery, New York, NY, USA,...
arXiv 2015
-
[65]
Ruihao Gong, Shihao Bai, Siyu Wu, Yunqian Fan, Zaijun Wang, Xiuhong Li, Hailong Yang, and Xianglong Liu. 2025. Past-Future Scheduler for LLM Serving under SLA Guarantees. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Rotterdam, Netherlands)(ASPLOS ’25). Association...
arXiv 2025
-
[66]
Yufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang, Xavier Servot, Onur Mutlu, Ravi Iyer, and Reetuparna Das. 2025. PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Rotterdam, Net...
arXiv 2025
-
[67]
Yue Guan, Xinwei Qiang, Zaifeng Pan, Daniels Johnson, Yuanwei Fang, Keren Zhou, Yuke Wang, Wanlu Li, Yufei Ding, and Adnan Aziz. 2025. Mercury: Unlocking Multi-GPU Operator Optimization for LLMs via Remote Memory Scheduling. InProceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles(Lotte Hotel World, Seoul, Republic of Korea)(SOSP ’25...
arXiv 2025
-
[68]
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving DNNs like clockwork: performance predictability from the bottom up. InProceedings of the 14th USENIX Conference on Operating Systems Design and Implementation (OSDI’20). USENIX Association, USA, Article 25, 20 pages
2020
-
[69]
Vincenzo Gulisano, Ricardo Jiménez-Peris, Marta Patiño-Martínez, Claudio Soriente, and Patrick Valduriez. 2012. StreamCloud: An Elastic and Scalable Data Streaming System.IEEE Transactions on Parallel and Distributed Systems23, 12 (2012), 2351–2365. doi:10.1109/TPDS.2012.24
-
[70]
Cong Guo, Rui Zhang, Jiale Xu, Jingwen Leng, Zihan Liu, Ziyu Huang, Minyi Guo, Hao Wu, Shouren Zhao, Junping Zhao, and Ke Zhang. 2024. GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Langu...
arXiv 2024
-
[71]
Rentong Guo, Xiaofan Luan, Long Xiang, Xiao Yan, Xiaomeng Yi, Jigao Luo, Qianya Cheng, Weizhi Xu, Jiarui Luo, Frank Liu, Zhenshan Cao, Yanliang Qiao, Ting Wang, Bo Tang, and Charles Xie. 2022. Manu: a cloud native vector database management system.Proc. VLDB Endow.15, 12 (Aug. 2022), 3548–3561. doi:10.14778/3554821.3554843
arXiv 2022
-
[73]
Yipin Guo and Siddharth Joshi. 2026. SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving. (2026). doi:10.48550/ARXIV. 2605.01708
-
[74]
Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024. HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models. InAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Associates, Inc., 59532–59569. doi:...
-
[75]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning (ICML’20). JMLR.org, Article 368, 10 pages
2020
-
[76]
Hayes, Kushal Kafle, Robik Shrestha, Manoj Acharya, and Christopher Kanan
Tyler L. Hayes, Kushal Kafle, Robik Shrestha, Manoj Acharya, and Christopher Kanan. 2020. REMIND Your Neural Network to Prevent Catastrophic Forgetting. InComputer Vision – ECCV 2020, Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, Cham, 466–483
2020
-
[77]
Congjie He, Yeqi Huang, Pei Mu, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang, and Luo Mai. 2025. WaferLLM: large language model inference at wafer scale. InProceedings of the 19th USENIX Conference on Operating Systems Design and Implementation(Boston, MA, USA)(OSDI ’25). USENIX Association, USA, Article 15, 17 pages
2025
-
[78]
Yongjun He, Haofeng Yang, Yao Lu, Ana Klimović, and Gustavo Alonso. 2025. Resource multiplexing in tuning and serving large language models. InProceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference(Boston, MA, USA)(USENIX ATC ’25). USENIX Association, USA, Article 97, 17 pages
2025
-
[79]
Martin Hirzel, Robert Soulé, Scott Schneider, Buğra Gedik, and Robert Grimm. 2014. A catalog of stream processing optimizations.ACM Comput. Surv.46, 4, Article 46 (March 2014), 34 pages. doi:10.1145/2528412
doi:10.1145/2528412 2014
-
[80]
Moritz Hoffmann, Andrea Lattuada, and Frank McSherry. 2019. Megaphone: latency-conscious state migration for distributed streaming dataflows. Proc. VLDB Endow.12, 9 (May 2019), 1002–1015. doi:10.14778/3329772.3329777
arXiv 2019
-
[81]
Connor Holmes, Masahiro Tanaka, Michael Wyatt, Ammar Ahmad Awan, Jeff Rasley, Samyam Rajbhandari, Reza Yazdani Aminabadi, Heyang Qin, Arash Bakhtiari, Lev Kurilenko, and Yuxiong He. 2024. DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference. (2024). doi:10.48550/ARXIV.2401.08671
-
[83]
Ke Hong, Xiuhong Li, Lufang Chen, Qiuli Mao, Guohao Dai, Xuefei Ning, Shengen Yan, Yun Liang, and Yu Wang. 2025. SOLA: Optimizing SLO Attainment for Large Language Model Serving with State-Aware Scheduling. InProceedings of Machine Learning and Systems, M. Zaharia, G. Joshi, and Y. Lin (Eds.), Vol. 7. MLSys. https://proceedings.mlsys.org/paper_files/paper...
2025
-
[85]
Chunyue Huang, Shuang Liu, Xinyi Zhang, Wenhao Li, Wei Lu, and Xiaoyong Du. 2025. Chimera: Mitigating Ownership Transfers in Multi-Primary Shared-Storage Cloud-Native Databases.Proc. VLDB Endow.18, 10 (June 2025), 3368–3381. doi:10.14778/3748191.3748201
arXiv 2025
-
[86]
Yibo Huang, Haowei Chen, Newton Ni, Yan Sun, Vijay Chidambaram, Dixin Tang, and Emmett Witchel. 2025. Tigon: a distributed database for a CXL pod. InProceedings of the 19th USENIX Conference on Operating Systems Design and Implementation(Boston, MA, USA)(OSDI ’25). USENIX Association, USA, Article 7, 20 pages
2025
-
[87]
Yuzhou Huang, Yapeng Jiang, Zicong Hong, Wuhui Chen, Bin Wang, Weixi Zhu, Yue Yu, and Zibin Zheng. 2025. Obscura: concealing recomputation overhead in training of large language models with bubble-filling pipeline transformation. InProceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference(Boston, MA, USA)(USENIX ATC ’25). USENIX Asso...
2025
-
[90]
Michael Isard, Mihai Budiu, Yuan Yu, Andrew Birrell, and Dennis Fetterly. 2007. Dryad: distributed data-parallel programs from sequential building blocks. InProceedings of the 2nd ACM SIGOPS/EuroSys European Conference on Computer Systems 2007(Lisbon, Portugal)(EuroSys ’07). Association for Computing Machinery, New York, NY, USA, 59–72. doi:10.1145/127299...
arXiv 2007
-
[91]
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: few-shot learning with retrieval augmented language models.J. Mach. Learn. Res.24, 1, Article 251 (Jan. 2023), 43 pages
2023
-
[92]
Sepehr Jalalian, Shaurya Patel, Milad Rezaei Hajidehi, Margo Seltzer, and Alexandra Fedorova. 2024. EXTMEM: enabling application-aware virtual memory management for data-intensive applications. InProceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference(Santa Clara, CA, USA)(USENIX ATC’24). USENIX Association, USA, Article 25, 12 pages
2024
-
[93]
Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park, Youngsok Kim, and Jinho Lee. 2024. Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System. In2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 345–360. doi:10.1109/HPCA57654.2024.00034
arXiv 2024
-
[94]
Keshav Vinayak Jha, Shweta Pandey, Murali Annavaram, and Arkaprava Basu. 2025. HyCache: hybrid caching for accelerating DNN input preprocessing pipelines. InProceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference(Boston, MA, USA)(USENIX ATC ’25). USENIX Association, USA, Article 26, 16 pages
2025
-
[96]
Sheng Jiang and Ming Liu. 2025. Building an Elastic Block Storage over EBOFs Using Shadow Views. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Philadelphia, PA, 1137–1153. https://www.usenix.org/conference/nsdi25/ presentation/jiang
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.