Pith. sign in

REVIEW 3 major objections 4 minor 70 references

Cloud abstractions for AI workloads

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read HarmonAIze argues that the root cause of AI workload inefficiency in multi-tenant clouds is the absence of abstractions for tenant–provider cooperation, and proposes a two-level abstraction design—micro-level, where providers execute…

desk verdict A clear, well-scoped vision paper for cooperative tenant-provider abstractions, but the load-bearing assumption that useful and non-sensitive infrastructure information can coexist is acknowledged, not solved. read the letter →

arxiv 2501.09562 v2 pith:652DNDVS submitted 2025-01-16 cs.DC

classification cs.DC
keywords cloudabstractionstenant-providercooperationcross-layeroptimizationAIworkloadscontrolloopscollectivecommunicationmulti-tenantworkloadintelligence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This position paper argues that AI workloads in multi-tenant clouds run inefficiently because tenants cannot see infrastructure state and providers cannot see workload intent. The proposed fix, HarmonAIze, is a layer of abstractions that splits control by timescale: providers execute micro-level loops over collective communication, batch size, and checkpointing using tenant-declared constraints, while tenants execute macro-level loops over scaling, parallelism, and hyperparameters using provider-supplied events. The paper argues that AI workloads' iterative, predictable, tightly synchronized character makes them the right first domain for such cooperation, and surveys seven optimization areas that would benefit. If the vision is right, shared control loops could turn one-off optimizations from research systems into standard cloud services.

What carries the argument

The central mechanism is a pair of abstraction levels. Micro-level abstractions keep API compatibility with existing AI libraries (such as collective communication interfaces) while letting tenants specify ranges and constraints, and they delegate data-path decisions to the provider, which can execute them in the hypervisor or on isolated infrastructure processing units. Macro-level abstractions give tenants a pub/sub interface to subscribe to infrastructure-level events (resource availability, failures) in aggregate or coarse-grained form, and allow multi-round negotiation between tenants and providers for strategic adaptations. The argument rests on the claim that AI workloads' predictability and tight synchronization make these cooperative control loops both feasible and high-value.

What would settle it

Run a controlled multi-tenant experiment comparing AI training jobs under HarmonAIze-style cooperative control (provider-side collectives and checkpointing, tenant-side scaling on provider events) against the same jobs on standard IaaS with static NCCL settings; if the cooperative cohort does not show measurably better job completion time, energy use, or failure recovery on realistic workloads with noisy neighbours, the core efficiency claim fails. A simpler disproof would be evidence that cloud providers cannot expose even aggregated infrastructure events at a granularity tenants can act on without leaking competitive information.

Watch

Extended reading notes

Core claim

HarmonAIze's central claim is that today's cloud fails AI workloads because tenants lack infrastructure insight and providers lack workload insight, and that the fix is a new layer of abstractions that splits control by timescale: providers execute micro-level control loops (algorithm selection, batch-size adjustment, checkpoint placement) using tenant-declared constraints, while tenants execute macro-level control loops (scaling, parallelization strategy, hyperparameters) using provider-supplied aggregated events. The paper argues this division matches who has the timely information and who has the strategic view, and surveys seven optimization opportunities (O1–O7) where cross-layer cooperation would amplify known gains. HarmonAIze deliberately focuses on AI workloads first because their iterative, accelerator-driven, and tightly synchronized nature makes them more predictable than general distributed systems, making the optimization problem tractable.

Load-bearing premise

Tenants and providers must both be willing to reveal enough—tenants their workload requirements, providers their infrastructure state—without leaking commercially sensitive details, and must trust each other enough to act on that shared information; if this mutual disclosure cannot happen, the cooperative control loops have nothing to feed on.

Editorial extensions

If this is right

  • Providers take over fine-grained, data-path decisions such as collective algorithm selection, batch-size adjustment, and checkpoint placement, using tenant-specified ranges as guardrails.
  • Tenants receive standardized, aggregate infrastructure events and can negotiate with providers, so scaling, parallelism, and hyperparameter choices can react to failures and congestion mid-run.
  • Known optimization gains—up to 2.4× collective communication speedup, 22.61× mean time between failures, and 53% energy reduction—would become attainable in public multi-tenant clouds rather than only in tightly coupled single-operator systems.
  • Adoption is incremental: tenants opt in by adding buy-in interfaces, while providers continue serving non-adopting tenants as today.
  • The first proof-of-concept can be assembled from existing open tools: scheduler simulators, distributed-job emulators, collectives-as-a-service, and runtime adaptation systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the premise holds, the same micro/macro split is a candidate for other synchronous, iterative workloads such as HPC simulations, even though the paper deliberately restricts itself to AI.
  • The buy-in deployment model implies a competitive dynamic the paper leaves implicit: early-adopting tenants should see measurable performance wins, which would push providers to standardize these interfaces to avoid churn.
  • A testable prediction following from the micro-level design is that provider-side execution of collectives on SmartNICs or IPUs outperforms tenant-side algorithm selection, because only the provider observes real-time topology and load.
  • The granularity of aggregate infrastructure events is the crux the paper flags but does not resolve: too coarse and tenants cannot act, too fine and providers leak sensitive information; finding the workable middle is an empirical question.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position paper proposes HarmonAIze, a set of micro- and macro-level cloud abstractions intended to enable cooperative optimization between AI workload tenants and cloud providers. The central idea is that tenant control loops should be informed by infrastructure-level insights while provider control loops are guided by workload-specific requirements, with micro-level abstractions supporting API-compatible, provider-executed fine-grained adaptations and macro-level abstractions providing tenant-executed strategic adaptations via interfaces such as pub/sub events. The paper identifies seven concrete optimization opportunities (O1--O7, in Section 4) and tabulates benefits from prior systems as representative of expected gains. It does not present an implementation or measurements; Section 5 outlines an incremental roadmap and briefly discusses adoption, timing, and open challenges including privacy.

Significance. If the HarmonAIze vision holds, it offers a useful organizing framework for research on tenant--provider cooperation in AI clouds, with a clear division of labor between micro-level provider control and macro-level tenant control. The paper's strengths are its concrete enumeration of seven optimization opportunities with up-to-date references, the explicit identification of building blocks for a prototype (e.g., Blox, MCCS, KungFu), and its honest treatment of adoption and standardization challenges. It is a genuine position paper: it frames testable hypotheses rather than claiming demonstrated results. The main significance is as a catalyst for discussion; the claimed benefits, however, are asserted rather than shown, and the privacy--utility tension is left unresolved.

major comments (3)
  1. [§3.1–3.2, §5] The central claim of Section 1 that HarmonAIze's two innovations are load-bearing requires an information channel that is simultaneously non-sensitive and actionable, but the paper never characterizes the granularity at which infrastructure insight remains actionable. Section 3.1 suggests aggregate or coarse-grained updates to safeguard sensitive details, and Section 3.2 reiterates that sensitive infrastructure details must not be exposed; however, for O2 (collective communication) and O3 (cluster scheduling), decisions such as algorithm selection and scaling require topology, link utilization, and other tenants' activity. Aggregating those signals to protect commercial confidentiality can remove precisely the information the tenant control loop needs. The benefits in Table 1 are from systems such as Pollux [43] and MCCS [56] that assume full visibility of both workload and infrastructure, so those numbers cannot be transferred without further argument to a regime with deliberately degraded information. This is a load-bearing gap; Section 5 lists privacy as open, but the authors should either provide a concrete privacy--utility analysis or moderate the claimed benefits accordingly.
  2. [§4, Table 1] Table 1 is introduced as 'select benefits that are representative of cooperative optimizations that we expect from HarmonAIze' (Section 5), but the cited systems are not instances of the HarmonAIze abstractions and in some cases require precisely the full visibility that HarmonAIze's coarse-grained interfaces deliberately withhold. For example, Pollux [43] co-adapts scheduling and batch size with direct control of cluster resources, and MCCS [56] offloads collectives as a service with infrastructure-level knowledge. As presented, the table conflates the headroom that cross-layer optimization could offer with the gains that HarmonAIze specifically would deliver. The paper should either relabel the table as evidence of headroom, or add a caveat that these numbers assume information quality that HarmonAIze's privacy constraints may not preserve.
  3. [§3.1, macro-level abstractions] The proposed pub/sub interface for macro-level updates mentions several possible mechanisms for resolving contention—priority-based allocation, reservation windows, or price-based auctioning—as though they were interchangeable (Section 3.1). These mechanisms impose very different incentive structures and information requirements: auctioning can elicit strategic behavior and reveal private valuations, reservation windows can reduce utilization, and priority rules require a notion of fairness across tenants. The paper's later claim of a 'natural alignment of incentives' (Section 5) therefore needs substantially more support; without it, the cooperative control-loop model is underspecified. At minimum, the authors should indicate which mechanism they envisage for the initial prototype described in Section 5.
minor comments (4)
  1. [O6] The text '22.61×%' in Section 4 (O6) contains a stray percentage sign; it should read '22.61×' or 'by a factor of 22.61'.
  2. [Figure 1] In Figure 1, 'MACRO-level' appears in all caps while 'micro-level' is lowercase; use consistent casing for the two levels.
  3. [Throughout] The abbreviation 'c.f.' (e.g., in Section 3) should be 'cf.' (confer), and 'check pointing' (Section 4, O6) should be 'checkpointing' for consistency.
  4. [Section 3.2] The phrase 'as much as a $30K [42]' in Section 4 (O6) is awkward; consider 'up to $30K [42]' or 'as much as $30K [42]'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a vision/position paper whose central claims are not derived from fitted parameters or self-cited uniqueness results.

full rationale

HarmonAIze is a position paper proposing new cloud abstractions; it contains no equations, no fitted parameters, and no derivation chain that reduces a prediction to its inputs. The claimed benefits in Table 1 are explicitly attributed to external literature results (e.g., KungFu, Parcae, Pollux, MCCS), and the paper frames HarmonAIze as an enabler that may amplify those benefits, not as a system that reproduces them. Some cited works have overlapping authors with the present paper (e.g., KungFu, Workload Intelligence, and a few others), but these citations are used as motivation, prior context, or roadmap components rather than as load-bearing justification for the core claim. The paper explicitly identifies open challenges such as privacy and adoption, and it does not assert that its abstractions have been validated. Because no step in the paper's argument is forced by definition, by construction, or by a self-citation chain, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The paper's design rests on assumptions about workload predictability, incentive alignment, and the feasibility of safe information sharing; none are empirically validated in the paper. No free parameters are fitted since no data is presented.

assumptions (4)
  • domain assumption AI workloads are iterative, accelerator-driven, and predictable at micro and macro timescales, making cooperative optimization tractable.
    Used to justify focusing on AI rather than general distributed systems; stated in Section 5 under 'Why not generalize beyond AI?'.
  • domain assumption Tenants and providers have aligned incentives to adopt cooperative abstractions and share information.
    Central to the adoption argument; discussed in Section 5 under 'Adoption' but not empirically validated.
  • domain assumption Infrastructure information can be shared with tenants in a form that is useful yet does not leak sensitive commercial details.
    Required for the micro and macro abstractions to work; the paper acknowledges the tension in Section 3.2 but provides no mechanism or proof.
  • domain assumption API-compatible abstractions can be introduced with minimal code changes to existing AI frameworks and libraries.
    Assumed in Section 3.1 ('drop-in module with minimal code changes'); no prototype demonstrates this.
invented entities (1)
  • HarmonAIze abstraction layer (micro and macro level interfaces)
    purpose: To enable cooperative optimization between tenants and providers for AI workloads.
    Proposed as a design; no implementation, measurements, or independently verifiable handle is provided in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cloud abstractions for AI workloads." pith.science (2026). https://pith.science/paper/652DNDVS

@misc{pith2026250109562,
  author       = {Pith},
  title        = {Pith review of: Cloud abstractions for AI workloads},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/652DNDVS}},
  note         = {Machine review of arXiv:2501.09562}
}
read the original abstract

AI workloads, often hosted in multi-tenant cloud environments, require vast computational resources but suffer inefficiencies due to limited tenant-provider coordination. Tenants lack infrastructure insights, while providers lack workload details to optimize tasks like partitioning, scheduling, and fault tolerance. We propose HarmonAIze to redefine cloud abstractions, enabling cooperative optimization for improved performance, efficiency, resiliency, and sustainability. We outline key opportunities and challenges this vision faces.

Figures

Figures reproduced from arXiv: 2501.09562 by the authors.

Figure 1
Figure 1. Architecture of HarmonAIze and overview of in￾teractions between tenant and provider to realize cross-layer optimization opportunities via micro- and MACRO-level cloud abstractions. Dashed call-outs indicate potential realization loci (e.g., at IPUs [27]) of optimized control loops. To address this gap, we propose widening traditional cloud abstractions to enable tenant-provider collaboration. Specifi￾cally, these a… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 63 canonical work pages

  1. [43]

    Ganger, and Eric P

    Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive Cluster Scheduling for Goodput- Optimized Deep Learning. InOSDI

  2. [56]

    Minjie Wang, Chien-chin Huang, and Jinyang Li. 2019. Supporting Very Large Models using Automatic Dataflow Graph Partitioning. In EuroSys

  3. [1]

    Abdelmoniem and Marco Canini

    Ahmed M. Abdelmoniem and Marco Canini. 2021. DC2: Delay-aware Compression Control for Distributed Machine Learning. InINFOCOM

  4. [2]

    Saurabh Agarwal, Amar Phanishayee, and Shivaram Venkataraman

  5. [3]

    Amazon Web Services. 2024. AWS SageMaker. https://aws.amazon. com/sagemaker

  6. [4]

    Mohammadreza Bayatpour, Nick Sarkauskas, Hari Subramoni, Ja- hanzeb Maqbool Hashmi, and Dhabaleswar K. Panda. 2021. BluesMPI: Efficient MPI Non-blocking Alltoall Offloading Designs on Modern Cloud abstractions for AI workloads APSys ’25, October 12–13, 2025, Seoul, Republic of Korea BlueField Smart NICs. InHigh Performance Computing

  7. [5]

    Muhammad Bilal, Marco Canini, Rodrigo Fonseca, and Rodrigo Ro- drigues. 2023. With Great Freedom Comes Great Opportunity: Re- thinking Resource Allocation for Serverless Functions. InEuroSys

  8. [6]

    Weilin Cai, Le Qin, and Jiayi Huang. 2025. MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training. InASPLOS

Show all 70 references
  1. [7]

    Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu. 2022. TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism.IEEE Transactions on Parallel and Distributed Systems33, 8 (2022)

  2. [8]

    Jiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, and Ennan Zhai. 2024. Crux: GPU- Efficient Communication Scheduling for Deep Learning Training. In SIGCOMM

  3. [9]

    Shubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, and Srinidhi Viswanatha. 2020. Balancing Efficiency and Fairness in Heterogeneous GPU Clusters for Deep Learning. In EuroSys

  4. [10]

    Jingrong Chen, Hong Zhang, Wei Zhang, Liang Luo, Jeffrey Chase, Ion Stoica, and Danyang Zhuo. 2022. NetHint: White-Box Networking for Multi-Tenant Data Centers. InNSDI

  5. [11]

    Xiaoqi Chen, Shay Vargaftik, and Ran Ben Basat. 2024. When ML Train- ing Cuts Through Congestion: Just-in-Time Gradient Compression via Packet Trimming. InHotNets

  6. [12]

    Jae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng, Nikhil Bansal, and Mosharaf Chowdhury. 2024. Reducing Energy Bloat in Large Model Training. InSOSP

  7. [13]

    Bryce Cronkite-Ratcliff, Aran Bergman, Shay Vargaftik, Madhusud- han Ravi, Nick McKeown, Ittai Abraham, and Isaac Keslassy. 2016. Virtualized Congestion Control. InSIGCOMM

  8. [14]

    Michael Dalton, David Schultz, Jacob Adriaens, Ahsan Arefin, Anshu- man Gupta, Brian Fahs, Dima Rubinstein, Enrique Cauich Zermeno, Erik Rubow, James Alexander Docauer, Jesse Alpert, Jing Ai, Jon Olson, Kevin DeCabooter, Marc de Kruijf, Nan Hua, Nathan Lewis, Nikhil Kasinadhun...

  9. [15]

    Jiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi, Dahua Lin, Harry Xu, Minjia Zhang, and Zhihao Jia. 2024. Parcae: Proactive, Liveput- Optimized DNN Training on Preemptible Instances. InNSDI

  10. [16]

    Benjamin Frank, Ingmar Poese, Yin Lin, Georgios Smaragdakis, Anja Feldmann, Bruce Maggs, Jannis Rake, Steve Uhlig, and Rick Weber

  11. [17]

    Google Cloud. 2024. Google Vertex AI. https://cloud.google.com/ vertex-ai

  12. [18]

    Google Cloud. 2024. Optimize Training Performance with Reduction Server in Vertex AI. https://cloud.google.com/blog/topics/developers- practitioners/optimize-training-performance-reduction-server- vertex-ai

  13. [19]

    Richard Graham, George Bosilca, Yong Qin, Bradley Settlemyer, Gi- lad Shainer, Craig Stunkel, Geoffroy Vallee, Brody Williams, Gerardo Cisneros-Stoianowski, Sebastian Ohlmann, and Markus Rampp. 2024. Optimizing Application Performance with BlueField: Accelerating Large-Message...

  14. [20]

    Tongzhou Gu, Jiawei Fei, and Marco Canini. 2024. OmNICCL: Zero- cost Sparse AllReduce with Direct Cache Access and SmartNICs. In NAIC

  15. [21]

    Tanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vi- jeev, Bhargav Gulavani, Nipun Kwatra, Ramachandran Ramjee, and Muthian Sivathanu. 2024. Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training Failures. InEuroSys

  16. [22]

    Keqiang He, Eric Rozner, Kanak Agarwal, Yu (Jason) Gu, Wes Felter, John Carter, and Aditya Akella. 2016. AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks. InSIGCOMM

  17. [23]

    Tao He, Xue Li, Zhibin Wang, Kun Qian, Jingbo Xu, Wenyuan Yu, and Jingren Zhou. 2024. Unicron: Economizing Self-Healing LLM Training at Scale. arXiv:2401.00134 [cs.DC]

  18. [24]

    Joseph, Randy Katz, Scott Shenker, and Ion Stoica

    Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, An- thony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. InNSDI

  19. [25]

    Samuel Hsia, Alicia Golden, Bilge Acun, Newsha Ardalani, Zachary DeVito, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. 2024. MAD- Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems. InISCA

  20. [26]

    Lexiang Huang, Anjaly Parayil, Jue Zhang, Xiaoting Qin, Chetan Bansal, Jovan Stojkovic, Pantea Zardoshti, Pulkit Misra, Eli Cortez, Raphael Ghelman, Íñigo Goiri, Saravan Rajmohan, Jim Kleewein, Ro- drigo Fonseca, Timothy Zhu, and Ricardo Bianchini. 2024. Work- load Intelligenc...

  21. [27]

    Intel. 2023. Intel Infrastructure Processing Unit (Intel IPU) ASIC E2000. https://www.intel.com/content/www/us/en/products/details/ network-io/ipu/e2000-asic.html

  22. [28]

    Xianyan Jia, Le Jiang, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang, Xinyuan Li, Langshi Chen, Yong Li, Zhen Zheng, Xiaoyong Liu, and Wei Lin. 2022. Whale: Efficient Giant Model Training over Heteroge- neous GPUs. InUSENIX ATC

  23. [29]

    Junchen Jiang, Xi Liu, Vyas Sekar, Ion Stoica, and Hui Zhang. 2014. EONA: Experience-Oriented Network Architecture. InHotNets

  24. [30]

    Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi...

  25. [31]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InSOSP

  26. [32]

    Dacheng Li, Hongyi Wang, Eric Xing, and Hao Zhang. 2022. AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness. InNeurIPS

  27. [33]

    Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka, Stas Bekman, Olatunji Ruwase, and Minjia Zhang. 2024. Universal Check- pointing: Efficient and Flexible Checkpointing for Large Scale Dis- tributed Training. arXiv:2406.18820 [cs.DC]

  28. [34]

    Zhiqi Lin, Youshan Miao, Quanlu Zhang, Fan Yang, Yi Zhu, Cheng Li, Saeed Maleki, Xu Cao, Ning Shang, Yilei Yang, Weijiang Xu, Mao Yang, Lintao Zhang, and Lidong Zhou. 2024. nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training. InOSDI

  29. [35]

    Banruo Liu, Mubarak Adetunji Ojewale, Yuhan Ding, and Marco Canini

  30. [36]

    Xuting Liu, Behnaz Arzani, Siva Kesava Reddy Kakarla, Liangyu Zhao, Vincent Liu, Miguel Castro, Srikanth Kandula, and Luke Marshall

  31. [37]

    Luo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis, Andrei- Octavian Brabete, and Peter Pietzuch. 2020. KungFu: Making Training APSys ’25, October 12–13, 2025, Seoul, Republic of Korea Canini et al. in Distributed Machine Learning Adaptive. InOSDI

  32. [38]

    Towards a Flexible and High-Fidelity Approach to Distributed DNN Training Emulation. InAPSys

  33. [39]

    Mustafa Rafique, Franck Cap- pello, and Bogdan Nicolae

    Avinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cap- pello, and Bogdan Nicolae. 2024. DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models. InHPDC

  34. [40]

    InSIGCOMM

    Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem. InSIGCOMM

  35. [41]

    Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravis- hankar Iyer. 2024. Queue Management for SLO-Oriented Large Lan- guage Model Serving. InSoCC

  36. [42]

    Ilia Markov, Kaveh Alim, Elias Frantar, and Dan Alistarh. 2024. L- GreCo: Layerwise-adaptive Gradient Compression For Efficient Data- parallel Deep Learning. InMLSys

  37. [44]

    NVIDIA. 2024. NVIDIA Collective Communication Library (NCCL). https://developer.nvidia.com/nccl

  38. [45]

    Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Cre- spo, and Dan Dennison

    D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Cre- spo, and Dan Dennison. 2015. Hidden Technical Debt in Machine Learning Systems. InNeurIPS

  39. [46]

    Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, Chao Wang, Peng Wang, Pengcheng Zhang, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, and Dennis Cai. 2024. Alibaba HPN: A Data Center Network for Large Languag...

  40. [47]

    Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2025. DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency. InHPCA

  41. [48]

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He

  42. [49]

    Yi Tay. 2024. Training Great LLMs Entirely from Ground Zero in the Wilderness. https://www.yitay.net/blog/training-great-llms-entirely- from-ground-zero-in-the-wilderness

  43. [50]

    Guanhua Wang, Olatunji Ruwase, Bing Xie, and Yuxiong He. 2024. FastPersist: Accelerating Model Checkpointing in Deep Learning. arXiv:2406.13768 [cs.DC]

  44. [51]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL]

  45. [52]

    Zhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang, Xinwei Fu, T. S. Eu- gene Ng, and Yida Wang. 2023. GEMINI: Fast Failure Recovery in Distributed Training with In-Memory Checkpoints. InSOSP

  46. [53]

    Tarnawski, Deepak Narayanan, and Amar Phanishayee

    Jakub M. Tarnawski, Deepak Narayanan, and Amar Phanishayee. 2021. Piper: Multidimensional Planner for DNN Parallelization. InNeurIPS

  47. [54]

    Carole-Jean Wu, Bilge Acun, Ramya Raghavendra, and Kim Hazelwood

  48. [55]

    Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, New- sha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien-Hsin...

  49. [57]

    Richard Yang, Arvind Krishnamurthy, Yanbin Grace Liu, and Abraham Silberschatz

    Haiyong Xie, Y. Richard Yang, Arvind Krishnamurthy, Yanbin Grace Liu, and Abraham Silberschatz. 2008. P4P: Provider Portal for Applica- tions. InSIGCOMM

  50. [58]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Syl- vain Gugger, Ma...

  51. [59]

    In EMNLP: System Demonstrations

    Transformers: State-of-the-Art Natural Language Processing. In EMNLP: System Demonstrations

  52. [60]

    Jie You, Jae-Won Chung, and Mosharaf Chowdhury. 2023. Zeus: Under- standing and Optimizing GPU Energy Consumption of DNN Training. InNSDI

  53. [61]

    Beyond Efficiency: Scaling AI Sustainably.IEEE Micro44, 5 (2024)

  54. [62]

    Xing, Joseph E

    Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Au- tomating Inter- and Intra-Operator Parallelism for Distributed Deep Learning. InOSDI

  55. [63]

    Yongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang, Ying Zhang, Matthew Lentz, and Danyang Zhuo. 2024. MCCS: A Service-based Approach to Collective Communication for Multi-Tenant Cloud. In SIGCOMM

  56. [65]

    Jihao Xin, Ivan Ilin, Shunkang Zhang, Marco Canini, and Peter Richtárik. 2023. Kimad: Adaptive Gradient Compression with Band- width Awareness. InDistributedML

  57. [66]

    Yifan Xiong, Yuting Jiang, Ziyue Yang, Lei Qu, Guoshuai Zhao, Shuguang Liu, Dong Zhong, Boris Pinzur, Jie Zhang, Yang Wang, Jithin Jose, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng, Yongqiang Xiong, and Lidong Zhou

  58. [67]

    InUSENIX ATC

    SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation. InUSENIX ATC

  59. [69]

    Yuxuan Zhao, Weikang Weng, Rob van Nieuwpoort, and Alexandru Uta. 2024. In Serverless, OS Scheduler Choice Costs Money: A Hybrid Scheduling Approach for Cheaper FaaS. InMiddleware

  60. [2013]

    Pushing CDN-ISP Collaboration to the Limit.SIGCOMM Comput. Commun. Rev.43, 3 (2013)

  61. [2020]

    ZeRO: Memory Optimizations Toward Training Trillion Param- eter Models. InSC

  62. [2024]

    InEu- roSys

    Blox: A Modular Toolkit for Deep Learning Schedulers. InEu- roSys

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.